6 Big Data
329
additional ability to evolve independently of the underlying RM layer in a much
more agile manner.
6.3 An Introduction to Big Data Modeling and Manipulation
With the evolution of computing technology, it is now possible to manage immense
volumes of data that previously could have only been handled at great expense by
supercomputers. Prices of computing and storage systems continue to drop, and as a
result, new techniques for distributed computing have become mainstream. The key
inflection point for Big Data occurred when companies like Yahoo!, Google, and
Facebook came to the realization that they there was an opportunity to monetize the
massive amounts of data collected. New technologies were required to create large
data stores, access those stores, and process huge amounts of data in near real-time.
The resulting solutions have transformed the data management market. In particular,
Hadoop, MapReduce, and Big Table proved to be the start of a new approach to data
management. These technologies address one of the most fundamental problems:
the need to process massive amounts of data efficiently, cost effectively, and in a
timely fashion.
6.3.1 Big Table
Big Table was developed by Google as a distributed storage system intended to
manage highly scalable structured data. Data is organized into tables with rows
and columns. Unlike a traditional relational database model, Big Table is a sparse,
distributed, persistent multidimensional sorted map. It is intended to store massive
volumes of data across a scalable array of commodity servers.
6.3.2 Pig
Pig is the high-level programming component running on top of Hadoop’s MapReduce component. Pig is a procedural language for creating MapReduce programs
used with Hadoop [12]. Pig was originally developed at Yahoo Research in 2006
to enable ad hoc execution of MapReduce jobs on very large data sets. The initial
language was called Pig Latin, enabling a variety of data manipulations in Hadoop.
It is a SQL-like language that enables a multi-query approach on a nested relational
data model where schema is optional. Apache Pig provides a rich set of built-in
operators to support data operations including joining, filtering, sorting, ordering,
nested data types, etc. on both structured and unstructured data. In addition, through
the User Defined Functions (UDF), Pig can invoke code in many other languages
Précédent

- 334/647

Suivant