322
N. Balac
was originally developed by Doug Cutting and Mike Cafarella in 2005 to support
distribution for the Nutch search engine project at Yahoo [7]. Doug named the
project after his then 2-year-old son’s toy elephant, hence the product’s current
pachyderm logo.
Hadoop allows applications based on MapReduce to run on large clusters of
commodity hardware. Hadoop is designed to parallelize data processing across
computing nodes to speed computations and minimize latency. Two major components of Hadoop are the massively scalable distributed file system that can support
petabytes of data and a massively scalable MapReduce engine that computes results
in batch mode.
Hadoop started out as a scalability solution to high volume and velocity batch
processing. The idea behind the processing paradigm called MapReduce [10] is to
provide a simple yet powerful computing framework that enables computations to
scale over big data easily. This type of system requires high efficiency. Therefore,
instead of moving data to computation, computation is moved closer to data in
Hadoop. Hadoop and MapReduce provide a shared and integrated foundation for
enabling this type of computation and seamlessly integrate additional tools to
support other necessary system functionalities. All of the modules in Hadoop are
designed with a fundamental assumption that hardware failures, whether of individual machines or racks of machines, are common and thus should be automatically
handled in software by the framework. The notion behind Hadoop’s reliability
requirements is based on the idea that if one computer fails once a year, then a
365-computer cluster will have a failure daily. If this number is scaled by an order
of magnitude, the cluster could be expected to have a hardware failure hourly. It is
essential for truly scalable systems to endure failure of any component.
Since data has been regarded as “new oil” or “new gold” in terms of an asset
value, a new approach to storing and computing has emerged. As the cost of storing
the data continues to decrease, organizations have developed a new approach of
keeping all data. Since data is growing rapidly in size and complexity, the “schema
on read style” has become the approach of choice. This approach enables all of the
data to be ingested in a rough form and then projected into the schema on the fly, as
it is pulled out of the stored location, thereby enabling experiments and new types
of analysis.
6.2.3 The Apache Hadoop Framework Components
All of the Hadoop modules are designed around a fundamental assumption of hardware failures. The entire Apache Hadoop “platform” is now commonly considered
to consist of a number of related projects including Apache Pig, Apache Hive,
Apache HBase, and others. These components will be described in the following
sections and are illustrated in Fig. 6.4.
N. Balac
was originally developed by Doug Cutting and Mike Cafarella in 2005 to support
distribution for the Nutch search engine project at Yahoo [7]. Doug named the
project after his then 2-year-old son’s toy elephant, hence the product’s current
pachyderm logo.
Hadoop allows applications based on MapReduce to run on large clusters of
commodity hardware. Hadoop is designed to parallelize data processing across
computing nodes to speed computations and minimize latency. Two major components of Hadoop are the massively scalable distributed file system that can support
petabytes of data and a massively scalable MapReduce engine that computes results
in batch mode.
Hadoop started out as a scalability solution to high volume and velocity batch
processing. The idea behind the processing paradigm called MapReduce [10] is to
provide a simple yet powerful computing framework that enables computations to
scale over big data easily. This type of system requires high efficiency. Therefore,
instead of moving data to computation, computation is moved closer to data in
Hadoop. Hadoop and MapReduce provide a shared and integrated foundation for
enabling this type of computation and seamlessly integrate additional tools to
support other necessary system functionalities. All of the modules in Hadoop are
designed with a fundamental assumption that hardware failures, whether of individual machines or racks of machines, are common and thus should be automatically
handled in software by the framework. The notion behind Hadoop’s reliability
requirements is based on the idea that if one computer fails once a year, then a
365-computer cluster will have a failure daily. If this number is scaled by an order
of magnitude, the cluster could be expected to have a hardware failure hourly. It is
essential for truly scalable systems to endure failure of any component.
Since data has been regarded as “new oil” or “new gold” in terms of an asset
value, a new approach to storing and computing has emerged. As the cost of storing
the data continues to decrease, organizations have developed a new approach of
keeping all data. Since data is growing rapidly in size and complexity, the “schema
on read style” has become the approach of choice. This approach enables all of the
data to be ingested in a rough form and then projected into the schema on the fly, as
it is pulled out of the stored location, thereby enabling experiments and new types
of analysis.
6.2.3 The Apache Hadoop Framework Components
All of the Hadoop modules are designed around a fundamental assumption of hardware failures. The entire Apache Hadoop “platform” is now commonly considered
to consist of a number of related projects including Apache Pig, Apache Hive,
Apache HBase, and others. These components will be described in the following
sections and are illustrated in Fig. 6.4.
