6 Big Data
343
Fig. 6.16 Spark’s internal job scheduling process
There are several advantages of DAG in Spark. In case of a lost RDD, Spark can
recover the information using the DAG and, with multiple levels of execution, can
execute a SQL query or ML operations with much more flexibility and efficiency
than MapReduce.
6.4.8 Spark SQL
Spark SQL is a Spark module for structured data processing on very large data sets.
Spark SQL provides Spark with additional information about the structure of data
and computation and uses this additional information to perform optimizations.
Spark SQL provides a fast execution engine by utilizing Spark as the underlying
execution engine for low-latency, interactive queries. It also provides the ability for
scale-out and failure recovery. The most common use of Spark SQL is to execute
SQL queries. However, it is also Hive compatible via Hive Query Language (HQL).
This allows it to read data from an existing Hive warehouse without a need to
change queries or move data [26]. Spark enables querying of various data sources
in addition to Hive tables including Parquet and JSON. In addition, Spark SQL
enables combining SQL queries with the data manipulations and complex analytics
supported by RDDs in Python, Java, and Scala [35].
Spark SQL provides three main capabilities for using structured and semistructured data. First, it provides a DataFrame abstraction in Python, Java, and
Scala, thereby simplifying the manipulation of structured datasets. Second, it can
read and write data in a variety of structured formats including JSON, Hive Tables,
and Parquet. Third, it enables data query using SQL. This can be accomplished both
inside a Spark program and from external tools that connect Spark SQL to thirdparty tools via standard database connectors like JDBC and ODBC [36].
Précédent

- 348/647

Suivant