6 Big Data
355
4. M. Cafarella, B. Lorica, D. Cutting, The next 10 years of Apache Hadoop. (O’Reilly Media,
2016). https://www.oreilly.com/ideas/the-next-10-years-of-apache-hadoop
5. G. Sanjay, G. Howard, L. Shun-Tak, The Google file system. SIGOPS Oper. Syst. Rev. 37(5),
29–43 (2003). https://doi.org/10.1145/1165389.945450
6. J. Dean, S. Ghemawat, MapReduce: Simplified data processing on large clusters. Commun.
ACM 51(1), 107–113 (2008). https://doi.org/10.1145/1327452.1327492
7. M. Bhandarkar, in 2010 IEEE International Symposium on Parallel & Distributed Processing
(IPDPS). MapReduce programming with apache Hadoop, (Atlanta, GA, 2010), pp. 1–1.
https://doi.org/10.1109/IPDPS.2010.5470377
8. P. Merla, Y. Liang, in 2017 IEEE International Conference on Big Data. Data analysis using Hadoop MapReduce environment, (Boston, MA, 2017), pp. 4783–4785.
https://doi.org/10.1109/BigData.2017.8258541
9. https://hadoop.apache.org/
10. D. Jeffrey, S. Ghemawat, in OSDI, MapReduce: Simplified data processing on large clusters
(2004)
11. S. Ghemawat, H. Gobioff, S. Leung, The Google file system, in Proceedings of the nineteenth
ACM symposium on Operating systems principles (SOSP ‘03), (ACM, New York, NY, USA,
2003), pp. 29–43. https://doi.org/10.1145/945445.945450
12. https://pig.apache.org/
13. https://hive.apache.org/
14. https://hadoop.apache.org/docs/r1.2.1/hdfs_design.html
15. K. Shvachko, H. Kuang, S. Radia, R. Chansler, in 2010 IEEE 26th Symposium on Mass Storage
Systems and Technologies (MSST), The Hadoop Distributed File System, (Incline Village, NV,
2010), pp. 1–10. https://doi.org/10.1109/MSST.2010.5496972
16. https://www.json.org/
17. https://avro.apache.org/
18. http://sqoop.apache.org/
19. E. Capriolo, D. Wampler, J. Rutherglen, Programming Hive: Data Warehouse and Query
Language for Hadoop, 1st edn. (O’Reilly Media, Sebastopol, CA, 2012). ISBN-13: 9781449319335. ISBN-10: 1449319335
20. K. Thulasiraman, M.N.S. Swamy, 5.7 Acyclic Directed Graphs, Graphs: Theory and Algorithms (Wiley, New York, 1992), p. 118. ISBN 978-0-471-51356-8
21. https://zookeeper.apache.org/
22. https://www.usenix.org/legacy/event/atc10/tech/full_papers/Hunt.pdf
23. H. Fang, in 2015 IEEE International Conference on Cyber Technology in Automation, Control,
and Intelligent Systems (CYBER), Managing data lakes in big data era: What’s a data lake and
why has it became popular in data management ecosystem, (Shenyang, 2015), pp. 820–824.
https://doi.org/10.1109/CYBER.2015.7288049
24. https://www.epic.com/
25. M. Zaharia, M. Chowdhury, M.J. Franklin, S. Shenker, I. Stoica, in Proceedings of the
2nd USENIX conference on Hot topics in cloud computing (HotCloud’10), Spark: Cluster
computing with working sets. (USENIX Association, Berkeley, CA, USA, 2010), pp. 10–10
26. https://spark.apache.org/
27. https://spark.apache.org/sql/
28. https://spark.apache.org/streaming/
29. https://spark.apache.org/mllib/
30. https://spark.apache.org/grapx
31. V.J. Srinivas, P. Srikanth, K. Thumati, S.H. Nallamala, in Proceedings International Journal
of Computer Science Trends and Technology (IJCST), A review study of Apache Spark in Big
Data processing, Vol. 4, Issue 3, May/Jun (2016)
32. M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M.J. Franklin, S. Shenker,
I. Stoica, in Proceedings of the 9th USENIX conference on Networked Systems Design and
Implementation (NSDI’12), Resilient distributed datasets: A fault-tolerant abstraction for inmemory cluster computing. (USENIX Association, Berkeley, CA, USA, 2012), pp. 2–2
355
4. M. Cafarella, B. Lorica, D. Cutting, The next 10 years of Apache Hadoop. (O’Reilly Media,
2016). https://www.oreilly.com/ideas/the-next-10-years-of-apache-hadoop
5. G. Sanjay, G. Howard, L. Shun-Tak, The Google file system. SIGOPS Oper. Syst. Rev. 37(5),
29–43 (2003). https://doi.org/10.1145/1165389.945450
6. J. Dean, S. Ghemawat, MapReduce: Simplified data processing on large clusters. Commun.
ACM 51(1), 107–113 (2008). https://doi.org/10.1145/1327452.1327492
7. M. Bhandarkar, in 2010 IEEE International Symposium on Parallel & Distributed Processing
(IPDPS). MapReduce programming with apache Hadoop, (Atlanta, GA, 2010), pp. 1–1.
https://doi.org/10.1109/IPDPS.2010.5470377
8. P. Merla, Y. Liang, in 2017 IEEE International Conference on Big Data. Data analysis using Hadoop MapReduce environment, (Boston, MA, 2017), pp. 4783–4785.
https://doi.org/10.1109/BigData.2017.8258541
9. https://hadoop.apache.org/
10. D. Jeffrey, S. Ghemawat, in OSDI, MapReduce: Simplified data processing on large clusters
(2004)
11. S. Ghemawat, H. Gobioff, S. Leung, The Google file system, in Proceedings of the nineteenth
ACM symposium on Operating systems principles (SOSP ‘03), (ACM, New York, NY, USA,
2003), pp. 29–43. https://doi.org/10.1145/945445.945450
12. https://pig.apache.org/
13. https://hive.apache.org/
14. https://hadoop.apache.org/docs/r1.2.1/hdfs_design.html
15. K. Shvachko, H. Kuang, S. Radia, R. Chansler, in 2010 IEEE 26th Symposium on Mass Storage
Systems and Technologies (MSST), The Hadoop Distributed File System, (Incline Village, NV,
2010), pp. 1–10. https://doi.org/10.1109/MSST.2010.5496972
16. https://www.json.org/
17. https://avro.apache.org/
18. http://sqoop.apache.org/
19. E. Capriolo, D. Wampler, J. Rutherglen, Programming Hive: Data Warehouse and Query
Language for Hadoop, 1st edn. (O’Reilly Media, Sebastopol, CA, 2012). ISBN-13: 9781449319335. ISBN-10: 1449319335
20. K. Thulasiraman, M.N.S. Swamy, 5.7 Acyclic Directed Graphs, Graphs: Theory and Algorithms (Wiley, New York, 1992), p. 118. ISBN 978-0-471-51356-8
21. https://zookeeper.apache.org/
22. https://www.usenix.org/legacy/event/atc10/tech/full_papers/Hunt.pdf
23. H. Fang, in 2015 IEEE International Conference on Cyber Technology in Automation, Control,
and Intelligent Systems (CYBER), Managing data lakes in big data era: What’s a data lake and
why has it became popular in data management ecosystem, (Shenyang, 2015), pp. 820–824.
https://doi.org/10.1109/CYBER.2015.7288049
24. https://www.epic.com/
25. M. Zaharia, M. Chowdhury, M.J. Franklin, S. Shenker, I. Stoica, in Proceedings of the
2nd USENIX conference on Hot topics in cloud computing (HotCloud’10), Spark: Cluster
computing with working sets. (USENIX Association, Berkeley, CA, USA, 2010), pp. 10–10
26. https://spark.apache.org/
27. https://spark.apache.org/sql/
28. https://spark.apache.org/streaming/
29. https://spark.apache.org/mllib/
30. https://spark.apache.org/grapx
31. V.J. Srinivas, P. Srikanth, K. Thumati, S.H. Nallamala, in Proceedings International Journal
of Computer Science Trends and Technology (IJCST), A review study of Apache Spark in Big
Data processing, Vol. 4, Issue 3, May/Jun (2016)
32. M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M.J. Franklin, S. Shenker,
I. Stoica, in Proceedings of the 9th USENIX conference on Networked Systems Design and
Implementation (NSDI’12), Resilient distributed datasets: A fault-tolerant abstraction for inmemory cluster computing. (USENIX Association, Berkeley, CA, USA, 2012), pp. 2–2
