336
N. Balac
Fig. 6.11 MapReduce
interactive execution
illustration
Disk
Query1
Query2
Query3
Query4
Result1
Results
Result3
Result4
Input
Data
File system Read
challenges. One aspect of the approach is to cache intermediate results in memory.
The second is to allow users to specify persistence in memory and partition the
dataset across nodes. In order to ensure fault tolerance, granular atomicity via
partitions and transaction logging were used instead of replication.
The highest-level unit of computation in MapReduce is a job that loads data,
applies a map function, shuffles, applies a reduce function, or writes data to
persistent storage. In Spark, on the other hand, the highest-level unit of computation
is an application that can be used for a single batch job, an interactive session with
multiple jobs, or a server repeatedly fulfilling requests. A Spark job can consist of
more than just a single map and reduce. Spark application processes can run on
its behalf even when it’s not running a job. Furthermore, multiple tasks can run
within the same executor, resulting in orders of magnitude faster performance when
compared to MapReduce.
Spark’s goal was to generalize MapReduce to support new applications and
enable more complex, often iterative or recursive, computations within same engine.
Two main additions to Hadoop’s approach were powerful enough to express these
types of computation and overcome deficiencies of the previous models: fast data
sharing and general Directed Acyclic Graphs (DAGs) for computation. These
approaches will be presented in more detail in the following sections. This approach
enables a much more efficient and much simpler approach, as end users can utilize
libraries instead of specialized systems to run complex computational workflows.
6.4.3 Resilient Distributed Datasets in Spark
The Resilient Distributed Dataset (RDD) is Spark’s fundamental data structure. It
is an immutable foundational distributed collection of objects [32]. The original
published paper proposed the concept of the RDD as a resilient, fault-tolerant
data structure. The RDD lineage graph is able to recompute missing or damaged
partitions due to node failures. RDDs are distributed with data residing on multiple
Précédent

- 341/647

Suivant