112
D. Lin et al.
Fig. 3. The Figure illustrates the latency with different number of mappers. Line with
‘o’ represents the weighted linear regression algorithm, and line with ‘+’ represents the
support vector machine algorithm.
5 Conclusion
In this paper, we establish a two-stage queueing model to mimic the MapReduce
process in consideration of the coupling effects between the map and reduce
stages. Also, with this model, we estimate the latency of data processing under
the MapReduce framework with multiple classification algorithms on various
datasets. The results show that our proposed algorithm in consideration of the
coupling effects outperforms the algorithms without considering the coupling
effects in the estimation of data processing latency. Also, the results show that
the bottleneck of data processing is usually at the mapping stage, and adding
mappers is more efficient than adding the reducers to reduce the latency of data
processing.
References
1. Grolinger K, Hayes M, Higashino WA, Heureux A, Allison DS, Capretz MAM (2014)
Challenges for MapReduce in big data. In: IEEE World Congress Services (SERVICES), pp 182–189
2. Najafabadi MM, Villanustre F, Khoshgoftaar TM, Seliya N, Wald R, Muharemagic
E (2015) Deep learning applications and challenges in big data analytics. J. Big
Data 2(1):1
3. Kune R, Konugurthi PK, Agarwal A, Chillarige RR, Buyya R (2016) The anatomy
of big data computing. Softw Pract Experience 46(1):79–105
4. Jagadish HV et al (2014) Big data and its technical challenges. Commun ACM
57(7):86–94
5. Lin D, Patrick J, Labeau F (2014) Estimating the waiting time of multi-priority
emergency patients with downstream blocking. Health Care Manage Sci 17(1):88–
99
6. Wissner-Gross A (2016) Datasets over algorithms. https://www.Edge.com.
Retrieved 8 Jan 2016
7. Sch¨ olkopf B, Platt J, Hofmann T (2007) Map-reduce for machine learning on multicore. In: Advances in neural information processing systems. Proceedings of the
2006 conference. MIT Press, pp 281–288
Précédent

- 124/679

Suivant