334
N. Balac
6.4.1 The Spark Ecosystem
While often compared to Hadoop and MapReduce, Spark is not a modified version
of Hadoop. Hadoop is simply one of several ways of implementing Spark. In fact,
Spark can run completely independently from Hadoop, powered by its own cluster
management.
Spark typically leverages Hadoop in two ways by utilizing its storage and
processing. As Spark has its own cluster management computation, it often times
only uses Hadoop for storage.
Spark Core is the underlying general execution engine for the Spark platform. It
enables in-memory computing and referencing datasets in external storage systems
(Fig. 6.9).
Spark SQL is a component on top of Spark Core that introduces a new data
abstraction called SchemaRDD, which provides support for structured and semistructured data [27].
Spark Streaming leverages Spark Core’s fast scheduling capability to perform
streaming analytics. It ingests data in mini-batches and performs RDD (Resilient
Distributed Datasets) transformations on those mini-batches of data [28].
Spark’s machine learning library called MLlib is a distributed machine learning
framework capable of taking advantage of the computational speedup related to
the distributed memory-based Spark architecture. Spark MLlib is often an order of
magnitude faster than Hadoop’s original, but now retired, Mahout library [29].
GraphX is a distributed graph-processing framework layered on top of Spark. It
provides an API for expressing graph computation. It also enables and optimizes
user-defined graphs and processing by leveraging Pregel abstraction API [30].
Spark provides built-in APIs, supports, and is compatible with many languages
and frameworks including Java, Scala, Python, R, Ruby, JavaScript, SparkSQL,
Hive, Pig, H20, etc.
Spark can run standalone or on top of a cluster computing framework such
as Hadoop. It handles batch, interactive, and real-time analysis within a single
framework. It provides native integration with Java, Python, and Scala, thereby
enabling programming at a higher level of abstraction.
One of the main advantages of Spark is that it is a general-purpose computing
engine that seamlessly encompasses data streaming management, data queries,
machine learning prediction, and real-time access to various analyses.
Spark
General Purpose ExecuƟon Engine
Spark SQL
Querying
Mllib
Machine Learning
Spark Streaming
Data Streaming
GraphX
Graph ComputaƟon
Fig. 6.9 Spark ecosystem
Précédent

- 339/647

Suivant