Zone-Based Resource Allocation Strategy
for Heterogeneous Spark Clusters
Yao Qin, Yu Tang
(B) , Xun Zhu, Chuanxiang Yan, Chenyao Wu, and Di Lin
University of Electronic Science and Technology of China, No. 4, Section 2,
North Jianshe Road, Chengdu, People’s Republic of China
yutang@uestc.edu.cn
Abstract. As a primary big data processing framework, Spark can
support memory computing to improve the computation efficiency. However, Spark cannot handle the situation of a heterogeneous cluster, in
which the nodes have different structures. Specifically, a primary problem in Spark is that the resource allocation strategy based on the number of homogeneous processor cores cannot adapt to the heterogeneous
cluster environment. To solve the above-mentioned problem, we propose
a zone-based resource allocation strategy based on heterogeneous Spark
cluster (ZbRAS) and implement such a strategy to improve the efficiency
of Spark. We compare the proposed strategy with the native resource
allocation strategy of Spark, and the comparison results show that our
proposed strategy can significantly enhance the execution speed of Spark
jobs in a heterogeneous cluster.
Keywords: Spark · Heterogeneity · Scheduling
1 Introduction
In recent years, Spark [1] and Hadoop [2] have emerged as an efficient framework
that is being extensively deployed to support a variety of big data applications.
And modern computer clusters supporting big data platforms such as Spark
are more and more heterogeneous with the continuous iterative updating of
the hardware in the cluster. The main challenge faced by researchers and IT
practitioners in optimizing Spark framework is how to make the best use of the
computation capacity in a heterogeneous cluster environment to deal with the
exponentially growing data volumes by an economically viable fashion.
Heterogeneous clusters mean that computers in a cluster have different hardware configurations, resulting in different computing performance of compute
nodes in the cluster. Therefore, scheduling optimization for heterogeneous clusters is an important direction in distributed computing. However, in heterogeneous clusters, the computing resource cannot be used reasonably, and the job
with higher priority cannot obtain higher performance computing resource.
c
Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 113–121, 2020.
https://doi.org/10.1007/978-981-15-0187-6_13
for Heterogeneous Spark Clusters
Yao Qin, Yu Tang
(B) , Xun Zhu, Chuanxiang Yan, Chenyao Wu, and Di Lin
University of Electronic Science and Technology of China, No. 4, Section 2,
North Jianshe Road, Chengdu, People’s Republic of China
yutang@uestc.edu.cn
Abstract. As a primary big data processing framework, Spark can
support memory computing to improve the computation efficiency. However, Spark cannot handle the situation of a heterogeneous cluster, in
which the nodes have different structures. Specifically, a primary problem in Spark is that the resource allocation strategy based on the number of homogeneous processor cores cannot adapt to the heterogeneous
cluster environment. To solve the above-mentioned problem, we propose
a zone-based resource allocation strategy based on heterogeneous Spark
cluster (ZbRAS) and implement such a strategy to improve the efficiency
of Spark. We compare the proposed strategy with the native resource
allocation strategy of Spark, and the comparison results show that our
proposed strategy can significantly enhance the execution speed of Spark
jobs in a heterogeneous cluster.
Keywords: Spark · Heterogeneity · Scheduling
1 Introduction
In recent years, Spark [1] and Hadoop [2] have emerged as an efficient framework
that is being extensively deployed to support a variety of big data applications.
And modern computer clusters supporting big data platforms such as Spark
are more and more heterogeneous with the continuous iterative updating of
the hardware in the cluster. The main challenge faced by researchers and IT
practitioners in optimizing Spark framework is how to make the best use of the
computation capacity in a heterogeneous cluster environment to deal with the
exponentially growing data volumes by an economically viable fashion.
Heterogeneous clusters mean that computers in a cluster have different hardware configurations, resulting in different computing performance of compute
nodes in the cluster. Therefore, scheduling optimization for heterogeneous clusters is an important direction in distributed computing. However, in heterogeneous clusters, the computing resource cannot be used reasonably, and the job
with higher priority cannot obtain higher performance computing resource.
c
Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 113–121, 2020.
https://doi.org/10.1007/978-981-15-0187-6_13
