6 Big Data
327
DATA
map()
map()
map()
reduce()
reduce()
O 1
O m
pairs
User Defined
Mapping FuncƟon Shuffle Phase
User Defined
Reduce FuncƟon
D 1
D 2
D n
Fig. 6.7 The MapReduce process
functions on a large amount of data in batch mode. The “map” function distributes
the programming tasks across a large number of commodity cluster nodes. It
handles the placement of the tasks in a way that balances the load and manages
recovery from failures. After the distributed computation is completed, another
function called “reduce” aggregates all the elements back together in a “shuffle”
and organizes the result. An example of MapReduce usage would be to determine a
word count across thousands of newspaper articles.
MapReduce works with the underlying file system and typically consists of
one JobTracker that receives the client’s MapReduce job requests (Fig. 6.8). The
JobTracker distributes processing to the available TaskTracker nodes in the cluster,
while striving to keep the work as close to the data as possible. JobTracker is aware
of which node contains the data and what neighboring processing is available. If
for some reason processing cannot be executed on the same data hosting node, then
priority is given to computing nodes in the same rack, thereby minimizing network
traffic.
If a TaskTracker fails or times out, that part of the job is rescheduled. The
TaskTracker on each node issues a separate Java Virtual Machine process to prevent
the TaskTracker from failing. A heartbeat is sent from the TaskTracker to the
JobTracker on a regular basis to check its status.
6.2.6 YARN
Apache Hadoop YARN is a sub-project of Hadoop. MapReduce underwent a
complete retrofit in an early version (v0.23) and the MapReduce 2.0 was designed
with YARN [18] as a key component. It separates the resource management and
327
DATA
map()
map()
map()
reduce()
reduce()
O 1
O m
User Defined
Mapping FuncƟon Shuffle Phase
User Defined
Reduce FuncƟon
D 1
D 2
D n
Fig. 6.7 The MapReduce process
functions on a large amount of data in batch mode. The “map” function distributes
the programming tasks across a large number of commodity cluster nodes. It
handles the placement of the tasks in a way that balances the load and manages
recovery from failures. After the distributed computation is completed, another
function called “reduce” aggregates all the elements back together in a “shuffle”
and organizes the result. An example of MapReduce usage would be to determine a
word count across thousands of newspaper articles.
MapReduce works with the underlying file system and typically consists of
one JobTracker that receives the client’s MapReduce job requests (Fig. 6.8). The
JobTracker distributes processing to the available TaskTracker nodes in the cluster,
while striving to keep the work as close to the data as possible. JobTracker is aware
of which node contains the data and what neighboring processing is available. If
for some reason processing cannot be executed on the same data hosting node, then
priority is given to computing nodes in the same rack, thereby minimizing network
traffic.
If a TaskTracker fails or times out, that part of the job is rescheduled. The
TaskTracker on each node issues a separate Java Virtual Machine process to prevent
the TaskTracker from failing. A heartbeat is sent from the TaskTracker to the
JobTracker on a regular basis to check its status.
6.2.6 YARN
Apache Hadoop YARN is a sub-project of Hadoop. MapReduce underwent a
complete retrofit in an early version (v0.23) and the MapReduce 2.0 was designed
with YARN [18] as a key component. It separates the resource management and
