data before being sent to the reducer is shuffled and sorted so that the reducer can
easily process it. Finally, the reducer performs an operation of reduction or
aggregation on the final data and this is followed by writing the final output to the
HDFS. The combiner does a similar task as the reducer but at the mapper lever
providing local lever aggregation or reduction.
III. YARN
Apache YARN stands for Yet Another Resource Negotiator. Before Hadoop 2.x,
the only framework which could run on Hadoop platform is MapReduce. The job
scheduling and resource negotiation is integrated with the MapReduce framework
and shared by Hadoop framework. The YARN provides the separate layer for job
scheduling and resource negotiation. It provides the platform for other programming framework like spark and storm, and many can run on Hadoop framework.
The basic architecture of YARN is shown in Fig. 6.
YARN has ResourceManager, NodeManager, Container, and ApplicationMaster.
Each container on datanode is specified with amount of CPU and memory, and it is
configurable. ResourceManager is run on namenode, and NodeManagers are run
on datanodes. Whenever a job is submitted, one container is allocated by a
ResourceManager on any datanode. This container process is called as
ApplicationMaster. This ApplicationMaster is responsible for all job management
and resource negotiation with ResourceManager. With the help of ResourceManager,
Fig. 6 Basic architecture of YARN showing various components
Turbo Analytics: Applications of Big Data and HPC in Drug …
355
easily process it. Finally, the reducer performs an operation of reduction or
aggregation on the final data and this is followed by writing the final output to the
HDFS. The combiner does a similar task as the reducer but at the mapper lever
providing local lever aggregation or reduction.
III. YARN
Apache YARN stands for Yet Another Resource Negotiator. Before Hadoop 2.x,
the only framework which could run on Hadoop platform is MapReduce. The job
scheduling and resource negotiation is integrated with the MapReduce framework
and shared by Hadoop framework. The YARN provides the separate layer for job
scheduling and resource negotiation. It provides the platform for other programming framework like spark and storm, and many can run on Hadoop framework.
The basic architecture of YARN is shown in Fig. 6.
YARN has ResourceManager, NodeManager, Container, and ApplicationMaster.
Each container on datanode is specified with amount of CPU and memory, and it is
configurable. ResourceManager is run on namenode, and NodeManagers are run
on datanodes. Whenever a job is submitted, one container is allocated by a
ResourceManager on any datanode. This container process is called as
ApplicationMaster. This ApplicationMaster is responsible for all job management
and resource negotiation with ResourceManager. With the help of ResourceManager,
Fig. 6 Basic architecture of YARN showing various components
Turbo Analytics: Applications of Big Data and HPC in Drug …
355
