120
Y. Qin et al.
Table 9. Result of Test 3
Native
Spark
ZbRAS Optimization
rate
Time 46.69 s
47.95 s
Node 4, 5, 6, 7 4, 5, 6,
7
Table 10. Result of Test 4
Native
Spark
ZbRAS Optimization
rate
Time 22.91 s
19.45 s 15.1%
Node 2, 4, 5, 6,
7, 8
1, 2, 3
strategy continues to start in zone 2, starting node 2 and node 3 for a total of
three Executors. The result of the final zone scheduling is much better than the
native Spark scheduling.
In conclusion, Test 1 and Test 2 reflect that ZbRAS has made full use of
the heterogeneity of the cluster and improved the execution speed of the job
by assigning high-performance computing resources to the job. Test 3 reflects
that ZbRAS can schedule the job to the specified area according to the user set
job priority. Test 4 reflects that ZbRAS will obtain computing resources in a
lower-level zone when the resources in the required zone are insufficient. Overall,
it shows that after zone scheduling, resource utilization is more reasonable, and
the execution speed of the application is increased by an average of 17%.
4 Conclusion
Overall, the experimental results show that, in heterogeneous Spark cluster, the
partition scheduling algorithm ZbRAS enables Spark to allocate resources based
on the heterogeneity of the cluster, gives the user the ability to allocate more
detailed resources for the job and makes the task scheduling more reasonable
and the task assignment more uniform, and the execution speed of Spark jobs
is significantly improved.
Acknowledgements. This work is jointly supported by the National Natural Science
Foundation of China (No. 61601082, No. 61471100, No. 61701503, No. 61750110527).
References
1. Armbrust M, Das T, Davidson A, Ghodsi A, Or A, Rosen J, Stoica I, Wendell P,
Xin R, Zaharia M (2015) Scaling spark in the real world: performance and usability.
Proc VLDB Endowment 8(12):1840–1843
2. Apache Hadoop (2011) http://hadoop.apache.org
3. Gao H, Yang Z, Bhimani J, Wang T, Wang J, Sheng B, Mi N (2017) Autopath:
harnessing parallel execution paths for efficient resource allocation in multi-stage
big data frameworks. In: 2017 26th international conference on computer communication and networks (ICCCN). IEEE, pp 1–9
4. Zaharia M, Chowdhury M, Franklin MJ, Shenker S, Stoica I (2010) Spark: cluster
computing with working sets. HotCloud 10(10–10):95
Y. Qin et al.
Table 9. Result of Test 3
Native
Spark
ZbRAS Optimization
rate
Time 46.69 s
47.95 s
Node 4, 5, 6, 7 4, 5, 6,
7
Table 10. Result of Test 4
Native
Spark
ZbRAS Optimization
rate
Time 22.91 s
19.45 s 15.1%
Node 2, 4, 5, 6,
7, 8
1, 2, 3
strategy continues to start in zone 2, starting node 2 and node 3 for a total of
three Executors. The result of the final zone scheduling is much better than the
native Spark scheduling.
In conclusion, Test 1 and Test 2 reflect that ZbRAS has made full use of
the heterogeneity of the cluster and improved the execution speed of the job
by assigning high-performance computing resources to the job. Test 3 reflects
that ZbRAS can schedule the job to the specified area according to the user set
job priority. Test 4 reflects that ZbRAS will obtain computing resources in a
lower-level zone when the resources in the required zone are insufficient. Overall,
it shows that after zone scheduling, resource utilization is more reasonable, and
the execution speed of the application is increased by an average of 17%.
4 Conclusion
Overall, the experimental results show that, in heterogeneous Spark cluster, the
partition scheduling algorithm ZbRAS enables Spark to allocate resources based
on the heterogeneity of the cluster, gives the user the ability to allocate more
detailed resources for the job and makes the task scheduling more reasonable
and the task assignment more uniform, and the execution speed of Spark jobs
is significantly improved.
Acknowledgements. This work is jointly supported by the National Natural Science
Foundation of China (No. 61601082, No. 61471100, No. 61701503, No. 61750110527).
References
1. Armbrust M, Das T, Davidson A, Ghodsi A, Or A, Rosen J, Stoica I, Wendell P,
Xin R, Zaharia M (2015) Scaling spark in the real world: performance and usability.
Proc VLDB Endowment 8(12):1840–1843
2. Apache Hadoop (2011) http://hadoop.apache.org
3. Gao H, Yang Z, Bhimani J, Wang T, Wang J, Sheng B, Mi N (2017) Autopath:
harnessing parallel execution paths for efficient resource allocation in multi-stage
big data frameworks. In: 2017 26th international conference on computer communication and networks (ICCCN). IEEE, pp 1–9
4. Zaharia M, Chowdhury M, Franklin MJ, Shenker S, Stoica I (2010) Spark: cluster
computing with working sets. HotCloud 10(10–10):95
