326
N. Balac
HDFS. It provides numerous benefits and has evolved into the de facto standard.
Avro’s lightweight and fast data serialization and deserialization enables fast
data ingestion. It stores metadata with the data itself and allows specification
of an independent schema for reading the files. It can quickly navigate to the
data collections in fast, random data access fashion. In addition, Avro files are
“splittable,” support block compression, and are accompanied by a wide and mature
set of open source tools.
The RC (Record Columnar) file format was the first columnar file in Hadoop.
It provides substantial compression and query performance benefits. However, it
does not support schema evaluation. Optimized RC Files (ORC) represent the
compressed version of RC files with additional improvements including enhanced
compression and faster querying.
The Parquet file format is another column-oriented data serialization standard
enabling compression, encodings, query performance benefits, and efficient data
analytics. This format has gained popularity as it became the choice of format for
Cloudera Impala. This optimization and usability contributed to its popularity in
other ecosystems as well.
Apache HBase is a scalable and distributed NoSQL database on HDFS for storing
key-value pairs. Keys are indexed, which typically enables fast access to the records.
The Apache Kudu file format is scalable and distributed table-based storage.
Kudu provides indexing and columnar data organization to achieve a balance
between ingestion speed and analytics performance. As in the case of HBase,
Kudu’s API enables modification of the data that is already stored in the system.
In general, three major factors should be considered when choosing the best
format for the task at hand: write performance, partial read performance, and
full read performance. These factors provide an indication of how fast the data
can be written, how fast individual columns can be read, and how fast can data
element be read from the data source. Columnar formats typically perform better
in terms of read performance. CSV and other non-compressed formats typically
demonstrate better write performance, but generally demonstrate slower reads due
to lack of compression. Some additional key factors that should be considered
while selecting the best file format include the type of the Hadoop distribution and
associated formats. Additionally, querying and processing requirements should be
considered alongside the processing tools. Further, extraction requirements merit
attention, especially when extracting data from a Hadoop environment into an
external database or other platforms. Finally, storage requirements are critical, as
volume may become a significant factor, and compression may be required.
6.2.5 MapReduce
The MapReduce paradigm is the core of the Hadoop system. MapReduce is
a distributed computing-based processing technique (Fig. 6.7). MapReduce was
designed by Google in order to satisfy the need for efficient execution of a set of
Précédent

- 331/647

Suivant