6 Big Data
341
using DAGs in Spark. DAG in Apache Spark is an alternative to the MapReduce. It is
a programming style used in distributed systems enabling multiple levels that form
a tree structure without having to write the intermediate results to disk. A DAG is a
finite directed graph with no directed cycles, consisting of finitely many vertices and
edges. Each edge is directed from one vertex to another in a consistently directed
sequence of edges that can never form a cycle [20].
The DAG concept has successfully been applied to Spark processing. A DAG is
represented by a set of vertices and edges, where vertices represent the RDDs and
edges represent the operation to be applied on the RDD. In order for action to be
executed on the RDD, Spark creates the DAG of the tasks to be executed. The DAG
is then submitted to the DAG scheduler that divides operators into stages of tasks in
order to perform parallel computation.
The principal unit of Spark’s computations is a job. It is typically a piece
of code that reads some input from HDFS, performs computation on the data,
and writes output data. Jobs are divided into stages. Stages are classified as a
Map or Reduce stages and are divided based on computational boundaries. Most
computations are executed over many stages. Each stage has some number of tasks,
and typically one task is executed on one partition of data on one machine. The
Executor is the process responsible for executing a task. The program/process
responsible for running the job over the Spark engine is typically referred to as a
driver. A driver runs on the master node, while the machine on which the executor
runs is referred to as the slave.
6.4.7 Spark Components
Spark applications run as independent sets of processes on a cluster, coordinated by
the SparkContext object located in the driver (see Fig. 6.14). The driver separates
the process to be executed and creates the SparkContext in order to schedule jobs
and negotiate with the cluster manager.
SparkContext can connect to several types of cluster managers (standalone Spark,
Mesos, YARN, or Kubernets), in order to allocate resources across applications.
Once connected, Spark acquires executors on nodes in the cluster. Executor
processes run computations and store application data. Subsequently, it sends
application code as defined by JAR or Python files passed to SparkContext, to the
executors. Finally, SparkContext sends tasks to the executors to run.
Spark jobs contain a series of operators and run on a set of data. All the operators
in a job are used to construct a DAG as shown in Fig. 6.15. The DAG is optimized by
rearranging and combining operators where possible. For example, if the submitted
Spark job contains a map operation followed by a filter operation, Spark’s DAG
optimizer will reorder the operators, as filtering reduces the number of records to
before applying the map operation.
Précédent

- 346/647

Suivant