4.2 Experimental Results and Analysis
Each set of tests has undergone multiple experiments and selected typical experimental
cases for analysis. The “average time of the task” is the average execution time of the
Task that takes the first stage of the job. This data is intended to illustrate the impact of
the execution time of a single task under different load conditions. “Maximum task
concurrency” refers to how many tasks are in parallel when the first stage is executed
on the specified computer. This data is intended to illustrate the task load on the node.
By analyzing the Task run and concurrency of the first stage, you can specifically
analyze the relationship between node load and Task execution.
Through experimental tests (Fig. 1), the load balancing scheduling operation time is
22.419 s, and the Spark scheduling operation time is 23.61 s. In test 1 (Fig. 1), the
initial load of all nodes is 0%. The native Spark scheduling and load balancing
scheduling in the above table yielded basically consistent runtime results. Although the
two nodes select different nodes, the initial load on each node is empty, and the time
performance of each running node is basically the same, indicating that the load
scheduling and the original Spark scheduling can achieve the same result without any
difference in load.
Table 2. Experimental test plan table
Test
number
Node
1 (%)
Node
2 (%)
Node
3 (%)
Node
4 (%)
Maximum task
concurrency
Processing data
volume (M)
1
0
0
0
0
12
384
2
0
0
90
0
12
384
3
0
0
0
0
9
2 8 8
4
9 0
3 0
2 0
9 0
9
2 8 8
Fig. 1. Experimental test 1
Task Scheduling Strategy for Heterogeneous Spark Clusters
135
Précédent

- 147/679

Suivant