330
N. Balac
including JRuby, Jython, and Java. This allows for the development of larger, more
complex applications.
Pig is typically used in ETL applications for describing how a process will extract
data from a source, transform it according to a rule set, and then load it into a
data store. Pig can ingest data from files, streams, or other sources using the UDF.
Once data is ingested, operations similar to the select command, various iterations,
and other complex transformations can be performed. Once the processing is
finalized, Pig stores the results of the transformations into the HDFS. Throughout
the processing steps, Pig scripts are translated into a series of MapReduce jobs
executed on the underlying Hadoop cluster.
6.3.3 Sqoop
Apache Sqoop is a tool designed for efficiently transferring bulk data between
Hadoop and structured data stores such as relational databases [18]. Sqoop is a
portmanteau that stands for SQL-to-Hadoop and is a simple command-line tool
with the several valuable capabilities. Sqoop has the capability to import individual
tables or entire databases to files in HDFS. It also generates Java classes to allow
interaction with imported data. Additionally, Sqoop provides the ability to import
from SQL databases directly into the Hive data warehouse within the Hadoop
environment, thereby enabling computing on the data very rapidly.
6.3.4 Hive
Hive is the data warehouse software platform that enables a SQL-like language for
facilitating, querying, and managing, large datasets residing in HDFS storage [13].
Often times referred to as the Hadoop data warehouse, Hive infrastructure sits on
top of Hadoop and provides data query and analysis. HiveQL is the mechanism used
to project structure onto the data and query the data using a SQL-like language [19].
HiveQL provides schema on read and transparently converts queries to MapReduce,
Apache Tez, and Spark jobs. All three execution engines can run in Hadoop
YARN. To accelerate queries, it provides indexes including bitmap indexes [13].
Additionally, Hive allows traditional and custom map and reduce mechanisms when
HiveQL might be insufficient.
Initially developed by Facebook, Apache Hive is now used and developed
industry wide. Hive supports analysis of large datasets stored in Hadoop’s HDFS
as well as compatible file systems such as the Amazon S3 filesystem.
N. Balac
including JRuby, Jython, and Java. This allows for the development of larger, more
complex applications.
Pig is typically used in ETL applications for describing how a process will extract
data from a source, transform it according to a rule set, and then load it into a
data store. Pig can ingest data from files, streams, or other sources using the UDF.
Once data is ingested, operations similar to the select command, various iterations,
and other complex transformations can be performed. Once the processing is
finalized, Pig stores the results of the transformations into the HDFS. Throughout
the processing steps, Pig scripts are translated into a series of MapReduce jobs
executed on the underlying Hadoop cluster.
6.3.3 Sqoop
Apache Sqoop is a tool designed for efficiently transferring bulk data between
Hadoop and structured data stores such as relational databases [18]. Sqoop is a
portmanteau that stands for SQL-to-Hadoop and is a simple command-line tool
with the several valuable capabilities. Sqoop has the capability to import individual
tables or entire databases to files in HDFS. It also generates Java classes to allow
interaction with imported data. Additionally, Sqoop provides the ability to import
from SQL databases directly into the Hive data warehouse within the Hadoop
environment, thereby enabling computing on the data very rapidly.
6.3.4 Hive
Hive is the data warehouse software platform that enables a SQL-like language for
facilitating, querying, and managing, large datasets residing in HDFS storage [13].
Often times referred to as the Hadoop data warehouse, Hive infrastructure sits on
top of Hadoop and provides data query and analysis. HiveQL is the mechanism used
to project structure onto the data and query the data using a SQL-like language [19].
HiveQL provides schema on read and transparently converts queries to MapReduce,
Apache Tez, and Spark jobs. All three execution engines can run in Hadoop
YARN. To accelerate queries, it provides indexes including bitmap indexes [13].
Additionally, Hive allows traditional and custom map and reduce mechanisms when
HiveQL might be insufficient.
Initially developed by Facebook, Apache Hive is now used and developed
industry wide. Hive supports analysis of large datasets stored in Hadoop’s HDFS
as well as compatible file systems such as the Amazon S3 filesystem.
