The How and Why of Fast Data Analytics with Apache Spark

The How and Why of Fast Data
Analytics with Apache Spark
with Justin Pihony  
@JustinPihony

Today’s agenda:
▪ Concerns
▪ Why Spark?
▪ Spark basics
▪ Common pitfalls
▪ We can help!
2

Concerns
▪ Am I too small?
4
▪ Will switching be too costly?
▪ Can I utilize my current infrastructure?
▪ Will I be able to find developers?
▪ Are there enough resources available?

object WordCount{
def main(args: Array[String])){
val conf = new SparkConf()
.setAppName("wordcount")
val sc = new SparkContext(conf)
sc.textFile(args(0))
.flatMap(_.split(" "))
.countByValue
.saveAsTextFile(args(1))
}
}
7
public class WordCount {
public static class Map extends Mapper<LongWritable, Text, Text, IntWritable> {
private final static IntWritable one = new IntWritable(1);
private Text word = new Text();
public void map(LongWritable key, Text value, Context context) throws IOException,
InterruptedException {
String line = value.toString();
StringTokenizer tokenizer = new StringTokenizer(line);
while (tokenizer.hasMoreTokens()) {
word.set(tokenizer.nextToken());
context.write(word, one);
}
}
}
public static class Reduce extends Reducer<Text, IntWritable, Text, IntWritable> {
public void reduce(Text key, Iterable<IntWritable> values, Context context)
throws IOException, InterruptedException {
int sum = 0;
for (IntWritable val : values) {
sum += val.get();
}
context.write(key, new IntWritable(sum));
}
}
public static void main(String[] args) throws Exception {
Configuration conf = new Configuration();
Job job = new Job(conf, "wordcount");
job.setOutputKeyClass(Text.class);
job.setOutputValueClass(IntWritable.class);
job.setMapperClass(Map.class);
job.setReducerClass(Reduce.class);
job.setInputFormatClass(TextInputFormat.class);
job.setOutputFormatClass(TextOutputFormat.class);
FileInputFormat.addInputPath(job, new Path(args[0]));
FileOutputFormat.setOutputPath(job, new Path(args[1]));
job.waitForCompletion(true);
}
}
Tiny CodeBig Code
Why Spark?

Why Spark?
8
Readability
Expressiveness
Fast
Testability
Interactive
Fault Tolerant
Unify Big Data

“Spark will kill MapReduce,
but save Hadoop.”
- http://insidebigdata.com/2015/12/08/big-data-industry-predictions-2016/

Big Data Unified API
13
Spark Core
Spark
SQL
Spark
Streaming
MLlib
(machine
learning)
GraphX
(graph)
DataFrames

Spark Mechanics
15
Worker WorkerWorker
Driver

Spark Mechanics
16
Spark Context
Worker WorkerWorker
Driver

Spark Context
17
Task creator
Scheduler
Data locality
Fault tolerance

RDD
18
▪ Resilient Distributed Dataset
▪ Transformations
- map
- filter
- …
▪ Actions
- collect
- count
- reduce
- …

Common Pitfalls
▪ Functional
▪ Out of memory
▪ Debugging
▪ …
21

Concerns
▪ Am I too small?
22
▪ Will switching from MapReduce be too costly?
▪ Can I utilize my current infrastructure?
▪ Will I be able to find developers?
▪ Are there enough resources available?

EXPERT SUPPORT
Why Contact Typesafe for Your Apache Spark Project?
Ignite your Spark project with 24/7 production SLA,
unlimited expert support and on-site training:
• Full application lifecycle support for Spark Core,
Spark SQL & Spark Streaming
• Deployment to Standalone, EC2, Mesos clusters
• Expert support from dedicated Spark team
• Optional 10-day “getting started” services
package
Typesafe is a partner with Databricks, Mesosphere
and IBM.
Learn more about on-site trainingCONTACT US

The How and Why of Fast Data Analytics with Apache Spark

More Related Content

What's hot(20)

Similar to The How and Why of Fast Data Analytics with Apache Spark(20)

More from Legacy Typesafe (now Lightbend)(14)

Recently uploaded(20)

The How and Why of Fast Data Analytics with Apache Spark