6 Big Data
337
nodes in a cluster. RDDs do not change once created and can only be transformed
using transformations to new RDDs.
Each dataset in a RDD is divided into logical partitions, which are often
computed on a variety of nodes of the cluster. By definition, RDD is a read-only,
partitioned, fault-tolerant collection of records that can be processed in parallel
manner [32]. RDDs can contain any type of Python, Java, Scala, or user-defined
classes and objects. RDDs can be created in one of two ways: either by parallelizing
an existing collection or referencing a dataset in an external storage system (HDFS,
HBase, etc.). Spark makes use of the concept of RDD to achieve faster and more
efficient MapReduce operations.
As shown in Fig. 6.12, iterative Operations on Spark RDDs store intermediate
results in a distributed memory and thereby offer much faster computation. In cases
when the distributed memory (RAM) is not sufficient to store intermediate results,
Spark will default to storing those results on the disk.
The interactive operations on Spark RDD illustrated in Fig. 6.13 are typically
utilized when different queries are executed on the same set of data repeatedly.
This particular data can be kept in memory to realize significant improvements in
execution times.
Each transformed RDD may be recomputed each time an action is executed on
that RDD. However, RDDs can also persist in memory, in which case Spark will
keep the elements on the cluster for considerably faster access the next time it is
queried. There is also support for persisting RDDs on disk in Spark and replication
across multiple nodes.
Disk
MapReduce
1
MapReduce
2
MapReduce
3
MapReduce
4
MapReduce
5
MapReduce
6
MapReduce
7
MapReduce
8
Disk
Input
Data
File System Read
Distributed Memory Write
Distributed Memory Read
File System Write
Distributed
memory
Fig. 6.12 Spark’s approach to fast data sharing for iterative operation
337
nodes in a cluster. RDDs do not change once created and can only be transformed
using transformations to new RDDs.
Each dataset in a RDD is divided into logical partitions, which are often
computed on a variety of nodes of the cluster. By definition, RDD is a read-only,
partitioned, fault-tolerant collection of records that can be processed in parallel
manner [32]. RDDs can contain any type of Python, Java, Scala, or user-defined
classes and objects. RDDs can be created in one of two ways: either by parallelizing
an existing collection or referencing a dataset in an external storage system (HDFS,
HBase, etc.). Spark makes use of the concept of RDD to achieve faster and more
efficient MapReduce operations.
As shown in Fig. 6.12, iterative Operations on Spark RDDs store intermediate
results in a distributed memory and thereby offer much faster computation. In cases
when the distributed memory (RAM) is not sufficient to store intermediate results,
Spark will default to storing those results on the disk.
The interactive operations on Spark RDD illustrated in Fig. 6.13 are typically
utilized when different queries are executed on the same set of data repeatedly.
This particular data can be kept in memory to realize significant improvements in
execution times.
Each transformed RDD may be recomputed each time an action is executed on
that RDD. However, RDDs can also persist in memory, in which case Spark will
keep the elements on the cluster for considerably faster access the next time it is
queried. There is also support for persisting RDDs on disk in Spark and replication
across multiple nodes.
Disk
MapReduce
1
MapReduce
2
MapReduce
3
MapReduce
4
MapReduce
5
MapReduce
6
MapReduce
7
MapReduce
8
Disk
Input
Data
File System Read
Distributed Memory Write
Distributed Memory Read
File System Write
Distributed
memory
Fig. 6.12 Spark’s approach to fast data sharing for iterative operation
