6 Big Data
331
6.3.5 HBase
Apache HBase is a column-oriented, distributed, and scalable database management
system that runs on top of HDFS. HBase is a key component of the Hadoop stack,
enabling fast, random access to large data sets. HBase is modeled after Google’s
BigTable, in order to handle massive data tables containing billions of rows and
millions of columns.
It is well suited for sparse data sets, which are common in numerous Big Data
use cases. Unlike relational database systems, HBase does not support a structured
query language like SQL. HBase applications are written in Java similar to a typical
MapReduce application. HBase also supports writing applications in Avro, REST,
and Thrift.
6.3.6 Oozie
Apache Oozie is a scalable, reliable, and extensible system workflow scheduler
system that manages and coordinates Apache Hadoop jobs while supporting
MapReduce, Pig, Hive, Sqoop, etc. Oozie workflow coordinator jobs are Directed
Acyclic Graphs (DAGs) of actions that are recurrent jobs triggered by time
frequency and data availability [20]. Oozie is integrated with the Hadoop stack,
typically with YARN, and supports numerous types of Hadoop jobs including Java
MapReduce, Pig, Hive, Streaming MapReduce, Sqoop, general-purpose Java code,
shell scripts, etc. Oozie itself is a Java Web application that combines multiple
jobs sequentially into one logical unit of work. Oozie Bundle enables packaging
of multiple coordinator and workflow jobs and management of the job’s life cycle.
It enables cluster administrators to develop complex data transformations with
multiple component tasks, thereby providing greater job control recurrence.
6.3.7 Zookeeper
Apache ZooKeeper [21] provides operational services for a Hadoop cluster by
enabling a distributed configuration service, a synchronization service, group
services, and a naming registry for distributed systems [22]. Distributed applications
use Zookeeper to store and mediate updates to important configuration information.
Due to the diversity of types of service implementations for applications,
management can become rather challenging when the applications are deployed.
ZooKeeper’s purpose is to extract the essence of these different services into a simple interface via a centralized coordination service. The service itself is distributed
and reliable supporting consensus, group management, and presence protocols.
Application-specific utilization consists of a mixture of specific components of
331
6.3.5 HBase
Apache HBase is a column-oriented, distributed, and scalable database management
system that runs on top of HDFS. HBase is a key component of the Hadoop stack,
enabling fast, random access to large data sets. HBase is modeled after Google’s
BigTable, in order to handle massive data tables containing billions of rows and
millions of columns.
It is well suited for sparse data sets, which are common in numerous Big Data
use cases. Unlike relational database systems, HBase does not support a structured
query language like SQL. HBase applications are written in Java similar to a typical
MapReduce application. HBase also supports writing applications in Avro, REST,
and Thrift.
6.3.6 Oozie
Apache Oozie is a scalable, reliable, and extensible system workflow scheduler
system that manages and coordinates Apache Hadoop jobs while supporting
MapReduce, Pig, Hive, Sqoop, etc. Oozie workflow coordinator jobs are Directed
Acyclic Graphs (DAGs) of actions that are recurrent jobs triggered by time
frequency and data availability [20]. Oozie is integrated with the Hadoop stack,
typically with YARN, and supports numerous types of Hadoop jobs including Java
MapReduce, Pig, Hive, Streaming MapReduce, Sqoop, general-purpose Java code,
shell scripts, etc. Oozie itself is a Java Web application that combines multiple
jobs sequentially into one logical unit of work. Oozie Bundle enables packaging
of multiple coordinator and workflow jobs and management of the job’s life cycle.
It enables cluster administrators to develop complex data transformations with
multiple component tasks, thereby providing greater job control recurrence.
6.3.7 Zookeeper
Apache ZooKeeper [21] provides operational services for a Hadoop cluster by
enabling a distributed configuration service, a synchronization service, group
services, and a naming registry for distributed systems [22]. Distributed applications
use Zookeeper to store and mediate updates to important configuration information.
Due to the diversity of types of service implementations for applications,
management can become rather challenging when the applications are deployed.
ZooKeeper’s purpose is to extract the essence of these different services into a simple interface via a centralized coordination service. The service itself is distributed
and reliable supporting consensus, group management, and presence protocols.
Application-specific utilization consists of a mixture of specific components of
