Through testing (Fig. 4), the load balancing scheduling operation time is 25.922 s,
the Spark scheduling operation time is 32.163 s, and the actual optimization percentage
is 19%. In Test 4, each node has different workloads. Load balancing scheduling
allocates different tasks on different nodes, and the result is better than native Spark
scheduling.
The performance of the above four tests shows that dynamic load scheduling can
avoid high-load nodes, allocate tasks according to the load on each node, and accelerate
the execution speed of the job.
5 Conclusion
In a production environment, Spark may run under complex load environments due to
resource isolation and resource reuse. Spark’s task scheduling is not based on node load
scheduling, which makes Spark unable to get the best performance under complex
cluster load conditions. Based on the above problem, the dynamic load-based Spark
task scheduling algorithm optimizes the task scheduling by using the dynamic load
status of each node of the cluster to improve the overall running speed of the operation.
The experimental results show that the dynamic load scheduling algorithm of heterogeneous Spark clusters gives users the ability to perform more detailed resource
allocation for jobs. The task scheduling is more reasonable, the task assignment is more
uniform, and the job execution speed is significantly improved.
Acknowledge. This work is jointly supported by the National Natural Science Foundation of
China (No. 61601082, No. 61471100, No. 61701503, No. 61750110527)
Fig. 4. Experimental test 4
Task Scheduling Strategy for Heterogeneous Spark Clusters
137
the Spark scheduling operation time is 32.163 s, and the actual optimization percentage
is 19%. In Test 4, each node has different workloads. Load balancing scheduling
allocates different tasks on different nodes, and the result is better than native Spark
scheduling.
The performance of the above four tests shows that dynamic load scheduling can
avoid high-load nodes, allocate tasks according to the load on each node, and accelerate
the execution speed of the job.
5 Conclusion
In a production environment, Spark may run under complex load environments due to
resource isolation and resource reuse. Spark’s task scheduling is not based on node load
scheduling, which makes Spark unable to get the best performance under complex
cluster load conditions. Based on the above problem, the dynamic load-based Spark
task scheduling algorithm optimizes the task scheduling by using the dynamic load
status of each node of the cluster to improve the overall running speed of the operation.
The experimental results show that the dynamic load scheduling algorithm of heterogeneous Spark clusters gives users the ability to perform more detailed resource
allocation for jobs. The task scheduling is more reasonable, the task assignment is more
uniform, and the job execution speed is significantly improved.
Acknowledge. This work is jointly supported by the National Natural Science Foundation of
China (No. 61601082, No. 61471100, No. 61701503, No. 61750110527)
Fig. 4. Experimental test 4
Task Scheduling Strategy for Heterogeneous Spark Clusters
137
