6 Big Data
323
Fig. 6.4 Apache Hadoop
basic components
HDFS
Hadoop Distributed File system
YARN
Resource Manager
MapReduce
Distributed Processing
Pig
ScripƟng
Hive
Query
Oozie
Workflow and
Scheduling
HBase
NoSQL Database
Zookeeper
CoordinaƟon
The Apache Hadoop framework is composed of the following modules:
• Hadoop Common: contains the libraries and utilities needed by other Hadoop
modules
• Hadoop Distributed File System (HDFS): a distributed file system that stores data
on the commodity machines, providing very high aggregate bandwidth across the
cluster
• Hadoop YARN: a resource-management platform responsible for managing compute resources in clusters and using them for scheduling of users’ applications
• Hadoop MapReduce: a programming model for large-scale data processing
Although the MapReduce Java code is common, any programming language
can be utilized to implement the “map” and “reduce” parts of the user’s program.
Apache Pig [12] and Apache Hive [13], among other related projects, expose higher
level user interfaces like Pig Latin and a SQL variant, respectively. The Hadoop
framework itself is mostly written in the Java programming language, with some
native code in C and command line utilities written as shell scripts.
The two primary components at the core of Apache Hadoop version 1 are
the Hadoop Distributed File System (HDFS) [14] and the MapReduce parallel
processing framework (Fig. 6.5). These are both open source projects, inspired
by technologies initially developed by Google. Hadoop’s MapReduce and HDFS
components originally derived from Google’s MapReduce and Google File System
(GFS) [11], respectively.
6.2.4 Hadoop Distributed File System
The Hadoop Distributed File System (HDFS) is a distributed, scalable, and portable
file system written in Java. Each node in a Hadoop instance typically has a single
NameNode, and a cluster of DataNodes forming the HDFS cluster [14]. Each
DataNode provides blocks of data using a HDFS-specific block protocol (Fig. 6.6).
Précédent

- 328/647

Suivant