Zone-Based Resource Allocation Strategy for Heterogeneous . . .
115
Spark, the heterogeneity of different computers is reflected in the time difference
between worker [4] in different computer performing the same Spark job. In
addition, since distributed computing requires data locality, the computers of
high data locality have the higher concurrency, the faster the job runs, and the
amount of computer concurrency is determined by the number of CPU cores. In
conclusion, the heterogeneity of computers in Spark is reflected in the difference
in the execution power of the CPU core and the number of CPU cores. Define
the heterogeneity of the computer in the Spark cluster by CPU cores (cpu cores)
and computing power (worker capacity), see Eq. 1.
Heterogeneity = (cpu cores, worker capacity)
(1)
Zone = (Heterogeneity,
m
i=1
W orker i)
(2)
Spark cluster = (Heterogeneity,
n
j=1
Zone j)
(3)
Then, define the zone as a grouping of computers with the same
Heterogeneity in the cluster, see Eq. 2. After the zone division, the Spark cluster
is composed of zones; see Eq. 3.
After the zone is introduced, in Spark, the task can be assigned to the specified zone to obtain the computing performance brought by the heterogeneity of
the zone according to the zone information. Figure 2 shows the task schedule in
Spark after the zone is introduced.
By dividing the cluster into zones, users can prioritize jobs based on their realtime requirements. Jobs with higher priorities run in high-performance zones,
and jobs with lower priorities run in low-performance zones. It means that cluster
resources have space for scheduling optimization, and scheduling optimization
for heterogeneity can improve the performance of Spark.
2.2 Zone Division Strategy
To enable Spark to allocate and schedule resources based on the hardware performance of each node in a heterogeneous cluster, the first problem to be solved
is how to make Spark aware of the heterogeneity of the cluster and obtain the
parameter worker capacity in Eq. 1. Computers are made up of complex hardwares. Different types of computers have a variety of hardware configurations,
so it is difficult to calculate the performance of various computers only through
hardware configurations. For the performance of the computer in the Spark
framework, we can test the results of the Spark task execution by executing
the Spark program and score the performance of the computer to obtain the
worker capacity of each worker in the Spark job.
The conventional method is to create a benchmark program. The benchmark
is a Spark job that runs only on one computer worker. We choose the classic
WordCount program, which reads text files and performs word frequency statistics on text data. The job result worker result is the execution time of the program, and time and speed are inversely proportional. Therefore, the reciprocal of
115
Spark, the heterogeneity of different computers is reflected in the time difference
between worker [4] in different computer performing the same Spark job. In
addition, since distributed computing requires data locality, the computers of
high data locality have the higher concurrency, the faster the job runs, and the
amount of computer concurrency is determined by the number of CPU cores. In
conclusion, the heterogeneity of computers in Spark is reflected in the difference
in the execution power of the CPU core and the number of CPU cores. Define
the heterogeneity of the computer in the Spark cluster by CPU cores (cpu cores)
and computing power (worker capacity), see Eq. 1.
Heterogeneity = (cpu cores, worker capacity)
(1)
Zone = (Heterogeneity,
m
i=1
W orker i)
(2)
Spark cluster = (Heterogeneity,
n
j=1
Zone j)
(3)
Then, define the zone as a grouping of computers with the same
Heterogeneity in the cluster, see Eq. 2. After the zone division, the Spark cluster
is composed of zones; see Eq. 3.
After the zone is introduced, in Spark, the task can be assigned to the specified zone to obtain the computing performance brought by the heterogeneity of
the zone according to the zone information. Figure 2 shows the task schedule in
Spark after the zone is introduced.
By dividing the cluster into zones, users can prioritize jobs based on their realtime requirements. Jobs with higher priorities run in high-performance zones,
and jobs with lower priorities run in low-performance zones. It means that cluster
resources have space for scheduling optimization, and scheduling optimization
for heterogeneity can improve the performance of Spark.
2.2 Zone Division Strategy
To enable Spark to allocate and schedule resources based on the hardware performance of each node in a heterogeneous cluster, the first problem to be solved
is how to make Spark aware of the heterogeneity of the cluster and obtain the
parameter worker capacity in Eq. 1. Computers are made up of complex hardwares. Different types of computers have a variety of hardware configurations,
so it is difficult to calculate the performance of various computers only through
hardware configurations. For the performance of the computer in the Spark
framework, we can test the results of the Spark task execution by executing
the Spark program and score the performance of the computer to obtain the
worker capacity of each worker in the Spark job.
The conventional method is to create a benchmark program. The benchmark
is a Spark job that runs only on one computer worker. We choose the classic
WordCount program, which reads text files and performs word frequency statistics on text data. The job result worker result is the execution time of the program, and time and speed are inversely proportional. Therefore, the reciprocal of
