utilization are usually selected to describe the load of the Spark program on the
computer. The computer load can be expressed as Eq. (1).
worker load ¼
ðcpu utilization þ memory utilization þ task competitionÞ
3
ð1Þ
Through the parameters, it can be calculated that the total time cpu_time of the
effective utilization of the CPU is Eq. (2).
cpu time ¼ user þ system þ nice þ idle þ iowait þ irq þ softirq
ð2Þ
For the calculation method, see Eq. (3).
cpu utilization ¼ 1 À
idle2 À idle1
cpu time2 À cpu time1
ð3Þ
How to calculate the memory utilization, see Eq. (4).
memory utilization ¼
mem total À mem free
mem total
ð4Þ
Whenever there is a Task that needs to be assigned an Executor, the minimum load
Executor in the load-ordered queue is selected to be dequeued. If the Executor resource
meets the running requirements of the Task, the Executor is assigned to the Task. The
load update method is shown in Eq. (5).
worker load ¼ 3 Â old load þ new load
ð
Þ =4
ð5Þ
The preset load in the schedule is shown in Eq. (6).
worker load ¼
spark:task:cpus
worker:cores
ð6Þ
3.2 Load Balancing Task Scheduling
Algorithm 1 describes dynamic load-based task scheduling in the form of pseudocode.
Algorithm 1 Dynamic load task scheduling algorithm
Input: node load queue load_worker_queue, to be scheduled task set TaskSets
Output: Dispatch Task to Executor Steps:
(1) updateLoad(load_worker_queue)
(2) for(worker in load_worker_queue){worker.load/=zone.cpu_capacity}
(3) sortedByLoad(load_worker_queue)
Task Scheduling Strategy for Heterogeneous Spark Clusters
133
computer. The computer load can be expressed as Eq. (1).
worker load ¼
ðcpu utilization þ memory utilization þ task competitionÞ
3
ð1Þ
Through the parameters, it can be calculated that the total time cpu_time of the
effective utilization of the CPU is Eq. (2).
cpu time ¼ user þ system þ nice þ idle þ iowait þ irq þ softirq
ð2Þ
For the calculation method, see Eq. (3).
cpu utilization ¼ 1 À
idle2 À idle1
cpu time2 À cpu time1
ð3Þ
How to calculate the memory utilization, see Eq. (4).
memory utilization ¼
mem total À mem free
mem total
ð4Þ
Whenever there is a Task that needs to be assigned an Executor, the minimum load
Executor in the load-ordered queue is selected to be dequeued. If the Executor resource
meets the running requirements of the Task, the Executor is assigned to the Task. The
load update method is shown in Eq. (5).
worker load ¼ 3 Â old load þ new load
ð
Þ =4
ð5Þ
The preset load in the schedule is shown in Eq. (6).
worker load ¼
spark:task:cpus
worker:cores
ð6Þ
3.2 Load Balancing Task Scheduling
Algorithm 1 describes dynamic load-based task scheduling in the form of pseudocode.
Algorithm 1 Dynamic load task scheduling algorithm
Input: node load queue load_worker_queue, to be scheduled task set TaskSets
Output: Dispatch Task to Executor Steps:
(1) updateLoad(load_worker_queue)
(2) for(worker in load_worker_queue){worker.load/=zone.cpu_capacity}
(3) sortedByLoad(load_worker_queue)
Task Scheduling Strategy for Heterogeneous Spark Clusters
133
