320
N. Balac
kinds of problems. Distributed processing typically segments large datasets, while
parallel processing simultaneously processes all data or subsets of data. In order
for modern systems to process Petabytes to Exabytes of data in a scalable and
practical manner, a system must be able to sustain partial failure. Any Big Data
processing system needs to continue processing in the face of failures without losing
any data and should be able to recover failed components and then allow them to
rejoin the process when ready. Additionally, failures during execution should not
affect the final result for consistency purposes, and the addition of resources should
automatically increase performance to enable seamless scalability. The Hadoop
framework is able to satisfy all of these requirements and has become one of the
leading components of the Big Data platforms for the last decade.
Hadoop is based on work done by Google in the early 2000s [4, 5]. More
specifically, Hadoop leverages the Google File System’s (GFS) MapReduce concepts, taking a fundamentally new approach to distributed computing [6–8]. Hadoop
meets the requirement for a low-cost, scalable, flexible, and fault-tolerant large-scale
system, while enabling the shared nothing architecture and the use of applications
written in high-level programming languages. The overall goal of Hadoop is to
execute computation where the data is stored, instead of moving large amount of
data to computational resources. Additionally, data is replicated multiple times on
the system for increased availability and reliability.
The Apache Hadoop software library [9] is an open source framework that
enables distributed processing of large data sets across clusters of computers using
simple programming models. It is designed to scale up from single servers to
thousands of machines, each offering local computation and storage. Rather than
relying on hardware to deliver high availability, the library itself is designed to detect
and handle failures at the application layer. This enables highly available services on
top of a commodity cluster where each machine may be prone to failures. Additional
open source projects have been built around the original Hadoop implementation
with the addition of Spark. Figure 6.3 depicts the basic Hadoop environment, which
enables large scale data management and computing.
6.2.1 Big Data System Architecture Components
This section examines in depth the key concepts required when considering implementation of Big Data, both from the technical and an analytics perspective. As
data systems became larger in recent years, performance became an acute concern.
An approach to efficiently solving a wide range of problems without needing to
change the underlying environment was needed. This type of approach would have
to take advantage of parallel processing of data and additional CPUs availability.
The goal of any analyst or data scientist, regardless of the field of study, is to work
with as much data as is available, build as many models as possible, and improve
model training time and accuracy by simply adding additional CPUs. The ultimate
goal of any Big Data system is to enable development of the best data analytics
N. Balac
kinds of problems. Distributed processing typically segments large datasets, while
parallel processing simultaneously processes all data or subsets of data. In order
for modern systems to process Petabytes to Exabytes of data in a scalable and
practical manner, a system must be able to sustain partial failure. Any Big Data
processing system needs to continue processing in the face of failures without losing
any data and should be able to recover failed components and then allow them to
rejoin the process when ready. Additionally, failures during execution should not
affect the final result for consistency purposes, and the addition of resources should
automatically increase performance to enable seamless scalability. The Hadoop
framework is able to satisfy all of these requirements and has become one of the
leading components of the Big Data platforms for the last decade.
Hadoop is based on work done by Google in the early 2000s [4, 5]. More
specifically, Hadoop leverages the Google File System’s (GFS) MapReduce concepts, taking a fundamentally new approach to distributed computing [6–8]. Hadoop
meets the requirement for a low-cost, scalable, flexible, and fault-tolerant large-scale
system, while enabling the shared nothing architecture and the use of applications
written in high-level programming languages. The overall goal of Hadoop is to
execute computation where the data is stored, instead of moving large amount of
data to computational resources. Additionally, data is replicated multiple times on
the system for increased availability and reliability.
The Apache Hadoop software library [9] is an open source framework that
enables distributed processing of large data sets across clusters of computers using
simple programming models. It is designed to scale up from single servers to
thousands of machines, each offering local computation and storage. Rather than
relying on hardware to deliver high availability, the library itself is designed to detect
and handle failures at the application layer. This enables highly available services on
top of a commodity cluster where each machine may be prone to failures. Additional
open source projects have been built around the original Hadoop implementation
with the addition of Spark. Figure 6.3 depicts the basic Hadoop environment, which
enables large scale data management and computing.
6.2.1 Big Data System Architecture Components
This section examines in depth the key concepts required when considering implementation of Big Data, both from the technical and an analytics perspective. As
data systems became larger in recent years, performance became an acute concern.
An approach to efficiently solving a wide range of problems without needing to
change the underlying environment was needed. This type of approach would have
to take advantage of parallel processing of data and additional CPUs availability.
The goal of any analyst or data scientist, regardless of the field of study, is to work
with as much data as is available, build as many models as possible, and improve
model training time and accuracy by simply adding additional CPUs. The ultimate
goal of any Big Data system is to enable development of the best data analytics
