Task Scheduling Strategy for Heterogeneous
Spark Clusters
Yu Liang, Yu Tang
(&) , Xun Zhu, Xiaoyuan Guo, Chenyao Wu,
and Di Lin
University of Electronic Science and Technology of China, Shahe Campus:
No. 4, Section 2, North Jianshe Road, 610054 Chengdu, Sichuan, People’s
Republic of China
yutang@uestc.edu.cn
Abstract. As a primary data processing and computing framework, Spark can
support memory computing, interactive computing, and querying in a huge
amount of data. Also, it can provide data mining, machine learning, stream
computing, and the other services. However, the strategy of allocating resources
among isomorphic processors cannot adapt to heterogeneous cluster environment due to its lack of load-based task scheduling. Therefore, we propose a
dynamic load scheduling algorithm for heterogeneous Spark clusters by regularly collecting load information from each of the cluster node. Such an algorithm can dramatically reduce the allocation of load to the nodes which are
already heavily loaded and in turn allocate more task to the idle nodes, thereby
speeding up the process of job allocation in Spark. The experimental results
show that the proposed algorithm can dramatically improve the computation
efficiency by dynamically loading among the nodes in a heterogeneous cluster.
Keywords: Spark platform Á Heterogeneous cluster Á Load balancing Á Task
scheduling
1 Introduction
The low-latency processing of massive data is a new technical challenge in the era of
big data. Spark [1] is a big data processing engine for in-memory computing, which
enables Spark to provide faster data processing. Spark performs lazy computing [2] by
manipulating elastic distributed datasets and adopts a simpler API to support more
flexible computing operations, demonstrating better performance than Hadoop [3] in
every respect.
Computers in a computer cluster have different hardware configurations, which
make them different in performance in Spark jobs. The development of cloud computing and the use of data centers make clusters more heterogeneous [4]. The rise of
machine learning has made clusters with computers with mixed CPU [5] and GPU [6]
architectures, making clusters heterogeneous.
Because the computing power of each node in a heterogeneous cluster [7] is
inconsistent, the same task allocation on different nodes will have different impact on
the node load. The default Spark task scheduling method does not consider cluster
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 131–138, 2020.
https://doi.org/10.1007/978-981-15-0187-6_15
Précédent

- 143/679

Suivant