Latency Estimation of Big Data Processing Under the MapReduce . . .
111
Table 2. Latency of various classification algorithms
Dataset Latency [s]
wlr
nb
log
svm
S
C
N
S
C
N
S
C
N
S
C
N
Am
20.7 20.3 19.2 19.8 19.6 18.3 20.2 19.8 18.6 18.9 18.6 17.4
Mo
6.1 6.0 4.9 5.6 5.5 4.5 5.9 5.8 4.8 5.0 5.0 4.0
Ya
3.1 3.0 2.4 2.8 2.7 2.2 2.9 2.9 2.3 2.5 2.4 1.7
Yo
0.6 0.6 0.4 0.5 0.5 0.3 0.6 0.6 0.3 0.5 0.5 0.3
Th
1.4 1.4 1.0 1.4 1.4 1.1 1.4 1.4 1.1 1.2 1.2 0.8
4.1 Latency of Various Classification Algorithms
The latency of various classification algorithms on various datasets is shown
in Table 2, in which ‘Am’ refers to the Amazon reviews dataset, ‘Mo’ refers
to MovieLens reviews dataset, ‘Ya’ refers to Yahoo music user ratings dataset,
‘Yo’ refers to Youtube comedy slam preference dataset, ‘Th’ refers to Thomson
Reuters text research collection dataset. Also, ‘wlr’ refers to the weighted linear
regression algorithm, ‘nb’ refers to the Naive Bayes algorithm, ‘log’ refers to
the logistic regression algorithm, and ‘svm’ refers to the support vector machine
algorithm. ‘S’ refers to the simulation results where we establish a cluster of
mappers and a cluster of reducers and compute the running time in the real lab
environment; ‘C’ refers to the numerical results in view of the coupling effects
(T
C
tot ); ‘N’ refers to the numerical results without considering the coupling effects,
i.e., the total latency of individual map and reduce stages (T
N
tot = T
M
W +T
M
S +T
R
S )
[7]. As shown in Table 2, the estimation by our algorithms in view of coupling
effects is closer to the latency in the simulation than the estimation without
considering coupling effects in [7].
4.2 Latency with Different Number of Mappers and Reducers
Table 2 shows the latency with only one mapper and one reducer. In the following, we investigate the latency with various numbers of mappers and reducers.
As shown in Fig. 3, the latency by the weighted linear regression algorithm and
the support vector machine algorithm almost linearly decrease with the number
of mappers and the number of reducers. Also, the latency can decrease to 1/10
when adding 10× mappers, while the latency is only around 1/3 when adding
10× reducers. Through adding more mappers, we can dramatically reduce the
latency. This result is in line with [7], which demonstrates that the processing
algorithms at the mappers have a higher computation complexity than those at
the reduces.
Précédent

- 123/679

Suivant