6 Big Data
335
6.4.2 The Core Difference Between Spark and Hadoop
MapReduce can enable users to write parallel computations using a set of highlevel operators without having to worry about work distribution and fault tolerance.
However, it also demonstrates a number of limitations for complex computational
tasks. While MapReduce is great at one-pass computation, it becomes rather
inefficient for multi-step algorithms due to the lack of efficient data sharing. Every
state between the map and reduce steps requires connection to the distributed file
system and is slow due to replication and disk storage. In most current frameworks,
the only way to reuse data between MapReduce jobs is to write data to an
external storage system and then read from the same location at later time. Since,
iterative and interactive applications require faster data sharing across parallel jobs,
Hadoop’s slow data sharing via MapReduce, due to replication, serialization, and
disk I/O can cause serious computational delays. This is a major contributor to the
often-presented statistic claiming that most of the Hadoop applications spend more
than 90% of the time on the HDFS read and write operations [31].
Iterative operations on MapReduce typically have a need to reuse intermediate
results across multiple computations in multi-stage applications. Figure 6.10 depicts
how a traditional Hadoop framework enables iterative operations via MapReduce,
showcasing data replication, disk I/O, and serialization overheads, causing the
overall computational time consequences.
Interactive operations on MapReduce are typically performed when ad hoc
queries are executed on the same subset of data. In this scenario, each query
will perform the disk I/O, which can dominate application execution time. Figure
6.11 illustrates how the traditional Hadoop framework accomplishes the interactive
queries on MapReduce.
As one might expect, large performance hits were experienced when complex
workflows were executed on large amounts of data. The need for executing complex workflows without writing intermediate results to disk after every operation
becomes a challenge. A new two-pronged approach was proposed to solve these
Disk
Map1
Map 2
Map 3
Map 4
Disk
Reduce1
Reduce2
Reduce3
Reduce4
Disk
Input
Data
File System Read
File System Write
File System Read
File System Write
Fig. 6.10 MapReduce execution of iterative operations
335
6.4.2 The Core Difference Between Spark and Hadoop
MapReduce can enable users to write parallel computations using a set of highlevel operators without having to worry about work distribution and fault tolerance.
However, it also demonstrates a number of limitations for complex computational
tasks. While MapReduce is great at one-pass computation, it becomes rather
inefficient for multi-step algorithms due to the lack of efficient data sharing. Every
state between the map and reduce steps requires connection to the distributed file
system and is slow due to replication and disk storage. In most current frameworks,
the only way to reuse data between MapReduce jobs is to write data to an
external storage system and then read from the same location at later time. Since,
iterative and interactive applications require faster data sharing across parallel jobs,
Hadoop’s slow data sharing via MapReduce, due to replication, serialization, and
disk I/O can cause serious computational delays. This is a major contributor to the
often-presented statistic claiming that most of the Hadoop applications spend more
than 90% of the time on the HDFS read and write operations [31].
Iterative operations on MapReduce typically have a need to reuse intermediate
results across multiple computations in multi-stage applications. Figure 6.10 depicts
how a traditional Hadoop framework enables iterative operations via MapReduce,
showcasing data replication, disk I/O, and serialization overheads, causing the
overall computational time consequences.
Interactive operations on MapReduce are typically performed when ad hoc
queries are executed on the same subset of data. In this scenario, each query
will perform the disk I/O, which can dominate application execution time. Figure
6.11 illustrates how the traditional Hadoop framework accomplishes the interactive
queries on MapReduce.
As one might expect, large performance hits were experienced when complex
workflows were executed on large amounts of data. The need for executing complex workflows without writing intermediate results to disk after every operation
becomes a challenge. A new two-pronged approach was proposed to solve these
Disk
Map1
Map 2
Map 3
Map 4
Disk
Reduce1
Reduce2
Reduce3
Reduce4
Disk
Input
Data
File System Read
File System Write
File System Read
File System Write
Fig. 6.10 MapReduce execution of iterative operations
