Latency Estimation of Big Data
Processing Under the MapReduce
Framework with Coupling Effects
Di Lin
(B) , Lingshuang Cai, Xiaofeng Zhang, Xiao Zhang, and Jiazhi Huo
School of Information and Software Engineering,
University of Electronic Science and Technology of China, Chengdu, China
lindi@uestc.edu.cn
Abstract. MapReduce is a model of processing large-scaled data with
parallel and distributed algorithms at a cluster, and it is composed of
two stages: a map stage for filtering and sorting data and a reduce stage
for the operation of summary. We develop a model with two connected
queues: one upstream queue for the data flow to access the mappers and
one downstream queue for the data flow to access the reducers. Also,
we analyze the latency of processing a large scale of data using queueing
models, in consideration of the coupling effects between these two queues
for map and reduce, respectively. Our analysis results on various datasets
and with various algorithms show that the MapReduce framework can
almost linearly speed up with increasingly more processors, and adding
mappers is usually more efficient than adding reducers to reduce the
latency when processing a large-scaled dataset.
Keywords: MapReduce · Coupling effects · Latency estimation · Big
data processing
1 Introduction
With the advances of mobile devices, sensing devices, social media, and Web
technologies, the amount of data is dramatically increasing at an unexpected
rate. In the era of big data, machine learning-based data analytics are viewed as
the primary driver of the big data revolution by extracting the insights of data.
A few algorithms have been adjusted to fit the machine learning algorithms for
large datasets, such as the algorithms adapted to the paradigm of MapReduce [1].
Quite a few studies have addressed the challenges and solutions in the design
of machine learning algorithms for big data processing under the MapReduce
framework. Big data process violates the assumption of loading the entire data
into memory or into a single file or disk at the training stage, and the operations
of several machine learning algorithms rely on this assumption. The violation of
c
Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 105–112, 2020.
https://doi.org/10.1007/978-981-15-0187-6_12
Processing Under the MapReduce
Framework with Coupling Effects
Di Lin
(B) , Lingshuang Cai, Xiaofeng Zhang, Xiao Zhang, and Jiazhi Huo
School of Information and Software Engineering,
University of Electronic Science and Technology of China, Chengdu, China
lindi@uestc.edu.cn
Abstract. MapReduce is a model of processing large-scaled data with
parallel and distributed algorithms at a cluster, and it is composed of
two stages: a map stage for filtering and sorting data and a reduce stage
for the operation of summary. We develop a model with two connected
queues: one upstream queue for the data flow to access the mappers and
one downstream queue for the data flow to access the reducers. Also,
we analyze the latency of processing a large scale of data using queueing
models, in consideration of the coupling effects between these two queues
for map and reduce, respectively. Our analysis results on various datasets
and with various algorithms show that the MapReduce framework can
almost linearly speed up with increasingly more processors, and adding
mappers is usually more efficient than adding reducers to reduce the
latency when processing a large-scaled dataset.
Keywords: MapReduce · Coupling effects · Latency estimation · Big
data processing
1 Introduction
With the advances of mobile devices, sensing devices, social media, and Web
technologies, the amount of data is dramatically increasing at an unexpected
rate. In the era of big data, machine learning-based data analytics are viewed as
the primary driver of the big data revolution by extracting the insights of data.
A few algorithms have been adjusted to fit the machine learning algorithms for
large datasets, such as the algorithms adapted to the paradigm of MapReduce [1].
Quite a few studies have addressed the challenges and solutions in the design
of machine learning algorithms for big data processing under the MapReduce
framework. Big data process violates the assumption of loading the entire data
into memory or into a single file or disk at the training stage, and the operations
of several machine learning algorithms rely on this assumption. The violation of
c
Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 105–112, 2020.
https://doi.org/10.1007/978-981-15-0187-6_12
