342
N. Balac
Cluster
Manager
Worker node 1
SparkContext
Driver Program
Cache
Task
Task
Task
Task
Worker Node 1
Worker node 1
Cache
Task
Task
Task
Task
Worker Node 2
Worker node 1
Cache
Task
Task
Task
Task
Worker Node 3
Fig. 6.14 Spark components
Fig. 6.15 DAG for a set of
Spark jobs
Start
Filter
Collect
Filter
Map
Reduce
Sort
Map
The Spark system is divided into various layers with individual responsibilities
in order to execute tasks efficiently. The layers are independent of each other. The
primary layer is the interpreter, and Spark uses a Scala interpreter. As commands
are entered in Spark console, Spark will create an operator graph. When an action
is executed, the graph is submitted to a DAG Scheduler. The DAG scheduler divides
the operator graph into a map and reduces stages. Each stage is comprised of
tasks based on partitions of the input data. The DAG scheduler orders operators
to optimize the graph, which is key to Spark’s fast performance. The final result
of a DAG scheduler is a set of stages that are passed to the Task Scheduler. The
Task Scheduler launches tasks via a cluster manager (Spark Standalone, Yarn,
Mesos, Kubernets). Nonetheless, it is unaware of any dependencies among stages
as illustrated in Fig. 6.16. The Worker executes the tasks; however, it knows only
about the code that is has received.
Précédent

- 347/647

Suivant