6 Big Data
325
The HDFS architecture includes a secondary NameNode; however its role is not
to take over when the primary NameNode goes offline. The secondary NameNode
connects with the primary NameNode on a regular basis and generates snapshots
of the primary NameNode’s directory information. This information is then saved
by the system and can be used to restart a failed primary NameNode in order to
create an up-to-date directory structure without having to repeat the entire set of file
system actions [15].
In the original Hadoop release, NameNode was the single metadata storage and
management information center. However, with the continually increasing number
of files, this approach created challenges. HDFS Federation solves the bottleneck
issue by allowing multiple name spaces served by distinct NameNodes and thereby
enables data awareness between the job tracker and task tracker. The job tracker
schedules map or reduces jobs, minimizing the amount of data movement. This can
have a significant impact on job completion times, especially with data-intensive
jobs.
6.2.4.1 Overview of Data Formats
There are a number of file and compression formats supported by the Hadoop
framework, each with a corresponding set of application-specific strengths and
weaknesses. HDFS enables several formats for storing data including HBase for
data access functionality and Hive for data management and querying functionality.
These file formats are designed for MapReduce or Spark computing engines for
specific purposes, ranging from basic analytics to machine learning applications.
Choosing the most appropriate file format can have a significant impact on
performance. It influences many aspects of the file system including read and write
times, the ability to split files into smaller components, enabling partial reads and
advanced compression support.
Hadoop enables the storage of text, binary, images, or other Hadoop-specific
formats. It provides built-in support for a number of formats specifically optimized
for Hadoop storage and processing. Some of the most common basic data formats
include text, CVS files, and JSON records. More complex formats such as Apache
Avro, Parquet, HBase, or Kudu can also be utilized. While text and CSV files are
very common, they do not support block compression and therefore often come with
a significant read performance cost. One common approach is to create a JSON
document in order to add structure to text files and utilize structured data in HDFS.
JSON is a common data format often used for asynchronous communication. JSON
stands for JavaScript Object Notation records and is an open-standard file format
that uses human-readable text to transmit data objects consisting of attribute-value
pairs and array data types. JSON files store metadata and can “split” files; however
it doesn’t support block compression [16].
Several additional more sophisticated and specialized file formats are available in
the Hadoop environment. One such format is Avro [17]. Avro is a data serialization
standard for the compact binary data format used for storing persistent data on
325
The HDFS architecture includes a secondary NameNode; however its role is not
to take over when the primary NameNode goes offline. The secondary NameNode
connects with the primary NameNode on a regular basis and generates snapshots
of the primary NameNode’s directory information. This information is then saved
by the system and can be used to restart a failed primary NameNode in order to
create an up-to-date directory structure without having to repeat the entire set of file
system actions [15].
In the original Hadoop release, NameNode was the single metadata storage and
management information center. However, with the continually increasing number
of files, this approach created challenges. HDFS Federation solves the bottleneck
issue by allowing multiple name spaces served by distinct NameNodes and thereby
enables data awareness between the job tracker and task tracker. The job tracker
schedules map or reduces jobs, minimizing the amount of data movement. This can
have a significant impact on job completion times, especially with data-intensive
jobs.
6.2.4.1 Overview of Data Formats
There are a number of file and compression formats supported by the Hadoop
framework, each with a corresponding set of application-specific strengths and
weaknesses. HDFS enables several formats for storing data including HBase for
data access functionality and Hive for data management and querying functionality.
These file formats are designed for MapReduce or Spark computing engines for
specific purposes, ranging from basic analytics to machine learning applications.
Choosing the most appropriate file format can have a significant impact on
performance. It influences many aspects of the file system including read and write
times, the ability to split files into smaller components, enabling partial reads and
advanced compression support.
Hadoop enables the storage of text, binary, images, or other Hadoop-specific
formats. It provides built-in support for a number of formats specifically optimized
for Hadoop storage and processing. Some of the most common basic data formats
include text, CVS files, and JSON records. More complex formats such as Apache
Avro, Parquet, HBase, or Kudu can also be utilized. While text and CSV files are
very common, they do not support block compression and therefore often come with
a significant read performance cost. One common approach is to create a JSON
document in order to add structure to text files and utilize structured data in HDFS.
JSON is a common data format often used for asynchronous communication. JSON
stands for JavaScript Object Notation records and is an open-standard file format
that uses human-readable text to transmit data objects consisting of attribute-value
pairs and array data types. JSON files store metadata and can “split” files; however
it doesn’t support block compression [16].
Several additional more sophisticated and specialized file formats are available in
the Hadoop environment. One such format is Avro [17]. Avro is a data serialization
standard for the compact binary data format used for storing persistent data on
