340
N. Balac
enabled by reusing intermediate in-memory results across multiple data-intensive
workloads, thereby avoiding movement of large amounts of data over the network.
6.4.5 Datasets and DataFrames in Spark
A Dataset is a distributed collection of data [33]. The Dataset is a new interface
added in Spark 1.6 that provides the benefits of RDDs including the strong
typing and powerful lambda functions, with the benefits of Spark SQL’s optimized
execution engine [33, 34]. A Dataset can be constructed from JVM objects and
then manipulated using functional transformations mentioned in the RDD section
including application of map, filter, etc. The Dataset API is available in Scala
and Java. Python and R do not have the support for the Dataset API, but due to those
languages’ dynamic nature, many of the benefits of the Dataset API are already
available and easily usable.
A DataFrame is a dataset organized into named columns. It is conceptually
equivalent to a table in a relational database or a data frame in R or Python, but with
richer built-in optimizations. DataFrames can be constructed from a wide array of
sources such as structured data files, tables in Hive, external databases, or existing
RDDs. The DataFrame API is available in Scala, Java, Python, and R, typically
represented by a Dataset of Rows [33]. More details on DataFrames are presented
below.
6.4.6 The Spark Processing Engine
While MapReduce is widely adopted for processing and generating large datasets
with a parallel, distributed algorithm on a cluster, more and more iterative and
interactive modes of operation have emerged that require faster data sharing across
parallel jobs. Data sharing is slow in MapReduce due to replication, serialization,
and disk IO. As noted earlier, studies have shown that most of the Hadoop
applications spend more than 90% of the time doing HDFS read-write operations
[25]. In contrast to Hadoop’s two-stage disk-based MapReduce paradigm, Spark’s
iterative in-memory approach enables added computing flexibility and significant
speedup. Additionally, Spark’s ability to load data into memory and query it
repeatedly enables scalable machine learning algorithm performance via the MLlib
library [29]. The ability to perform sophisticated, advanced analytics is one of the
main advantages of Spark, as it also supports SQL queries, streaming data, machine
learning (ML), and graph algorithms.
While MapReduce can achieve complex tasks by defining and chaining various
maps and reduce tasks, it is limited to one directional, sequential execution of the
mappers and reducers. This limitation has been overcome by allowing task definition
Précédent

- 345/647

Suivant