110
D. Lin et al.
P Q =
(K 1 ρ)
K1
K 1 !(1 − ρ)
K1−1
p=1
(K 1 ρ)
p
p!
+
∞
p=K1
(K 1 ρ)
p
K 1 !K
p−K1
1
−1
(5)
where ρ =
λ
μK1 .
Equation (5) indicates that T
M
W depends on μ, and Eq. (1) shows that the
computation of μ depends on mean(min
i
W i ), which will be detailed in the following Sect. 3.2.
3.2 Estimation of the Waiting Time to Be Transferred to Reducers
In the following, we study how to compute the average waiting time to be transferred to reducers, i.e., mean(min
i
W i ). Denote F i (z) as the cumulative distribution function (CDF) of W i (i = 1, . . . , K 2 ). Then the CDF of min
i
W i (denoted
as F min (z)) can be expressed as
F min (z) = 1 −
K2
i=1
[1 − F i (z)]
(6)
Based on Eq. (6), we can achieve the mean of min
i
W i as
mean(min
i
W i ) =
zdF min (z)
(7)
By substituting Eq. (6) into Eq. (7), we can compute the average waiting time
to be transferred to reducers, i.e., mean(min
i
W i ).
4 Simulation Results
In this section, we employ multiple classification algorithms under the MapReduce framework on the real data of a few public datasets [6], including the
Amazon reviews dataset, MovieLens reviews dataset, Yahoo music user ratings dataset, Youtube comedy slam preference dataset, and Thomson Reuters
text research collection dataset. The basic information of the above-mentioned
datasets is summarized in Table 1.
Table 1. Information of datasets for classification algorithms
Dataset
Samples
Amazon reviews
82,000,000
MovieLens reviews
22,000,000
Yahoo music user rating
10,000,000
Youtube comedy slam preference 1,138,562
Thomson Reuters text research
1,800,370
Précédent

- 122/679

Suivant