heterogeneity and lack of load-based task scheduling strategies. Therefore, this paper
proposes a load scheduling algorithm that can consider cluster node heterogeneity and
node resource dynamics.
2 Research Background
2.1 Spark Scheduling Strategy
The scheduling model in Spark refers to the process in which Spark writes the submitted jobs and data, how it is parsed, split and transformed by Spark, and how it is
distributed to each computing node in the cluster for distributed computing and results.
During the scheduling process, Spark’s scheduling components and models are
involved, as well as the job scheduling algorithms provided by Spark. Spark’s job
scheduling strategy is divided into two types: FIFO [8] (first in, first out) and FAIR (fair
scheduling).
2.2 Task Scheduling Process
When scheduling Task to Executor, if you can fully understand the load of Spark
cluster and introduce load balancing algorithm [8], you can make full use of low-load
nodes, make the load of nodes in the cluster more even. More rational use of cluster
computing resources can speed up the operation of the entire job.
In the above task scheduling process, there is no load-based scheduling process.
There are many factors that cause this situation. There are two main points, namely
resource reuse and container technology.
Whether a physical cluster deploys multiple computing architectures or uses container technology for resource isolation, it means that the operation of Spark in a
production environment may share computing resources with other computing services. The operational load of other computing frameworks will definitely affect the
Spark running tasks. If you can schedule tasks based on node load, Spark can still
maintain high-performance computing services in complex hardware and software
environments.
3 Spark Load Balancing Task Scheduling Strategy
3.1 Spark Load Definition
The process of Spark assigning Executors to Tasks is random. Introducing a load
balancing algorithm here will optimize this process and speed up the progress of the
job.
In order to sort according to the load, a calculation method of the node load is
needed to judge the load level of each node. Because Spark is a computationally
intensive and in-memory computing framework, the main computing resources used
are CPU and memory. The CPU utilization, system task queue length, and memory
132
Y. Liang et al.
proposes a load scheduling algorithm that can consider cluster node heterogeneity and
node resource dynamics.
2 Research Background
2.1 Spark Scheduling Strategy
The scheduling model in Spark refers to the process in which Spark writes the submitted jobs and data, how it is parsed, split and transformed by Spark, and how it is
distributed to each computing node in the cluster for distributed computing and results.
During the scheduling process, Spark’s scheduling components and models are
involved, as well as the job scheduling algorithms provided by Spark. Spark’s job
scheduling strategy is divided into two types: FIFO [8] (first in, first out) and FAIR (fair
scheduling).
2.2 Task Scheduling Process
When scheduling Task to Executor, if you can fully understand the load of Spark
cluster and introduce load balancing algorithm [8], you can make full use of low-load
nodes, make the load of nodes in the cluster more even. More rational use of cluster
computing resources can speed up the operation of the entire job.
In the above task scheduling process, there is no load-based scheduling process.
There are many factors that cause this situation. There are two main points, namely
resource reuse and container technology.
Whether a physical cluster deploys multiple computing architectures or uses container technology for resource isolation, it means that the operation of Spark in a
production environment may share computing resources with other computing services. The operational load of other computing frameworks will definitely affect the
Spark running tasks. If you can schedule tasks based on node load, Spark can still
maintain high-performance computing services in complex hardware and software
environments.
3 Spark Load Balancing Task Scheduling Strategy
3.1 Spark Load Definition
The process of Spark assigning Executors to Tasks is random. Introducing a load
balancing algorithm here will optimize this process and speed up the progress of the
job.
In order to sort according to the load, a calculation method of the node load is
needed to judge the load level of each node. Because Spark is a computationally
intensive and in-memory computing framework, the main computing resources used
are CPU and memory. The CPU utilization, system task queue length, and memory
132
Y. Liang et al.
