328
N. Balac
Fig. 6.8 MapReduce, detailed view
processing components enabling a broader array of interaction patterns for data
stored in HDFS. The YARN-based architecture provides a more general processing
platform that is not constrained by MapReduce limitations.
The fundamental concept underlying YARN is to split up the two major functionalities of the Job Tracker, resource management and job scheduling/monitoring, into
separate functionalities. The Global Resource Manager (RM) and per-application
Application Master (AM) are key components. The RM is the ultimate authority
that arbitrates resources among all the applications in the system. The AM is a
framework-specific library tasked with negotiating resources from the RM and
working with the Node Manager(s) to execute and monitor the tasks [18].
YARN enhances the power of a Hadoop-based cluster in several key ways. First,
it enables a higher level of scalability as the processing power in data centers continues to grow quickly. The YARN RM focuses solely on scheduling and therefore
is able to manage extremely large clusters very efficiently. Furthermore, YARN
enables compatibility with existing MapReduce applications without disruption to
the existing processes. Additionally, YARN significantly improves cluster utilization
by enabling workloads beyond that of MapReduce. The MapReduce RM is a pure
scheduler that optimizes cluster utilization according to specified criteria such as
capacity guarantees, fairness, and SLAs. In contrast, YARN enables additional
programming models for real-time processing such as Spark, graph processing,
machine learning, and iterative modeling. YARN’s processing approach has the
Précédent

- 333/647

Suivant