344
N. Balac
Fig. 6.17 Many ways to create a DataFrame in Spark
6.4.9 Spark DataFrames
A DataFrame in Spark represents a distributed collection of data organized into
named columns [33]. A DataFrame is conceptually equivalent to a table in a
relational database, a data frame in R or Python’s Panda DataFrame, but with
additional optimizations for the Spark engine. DataFrames support and can be
constructed from a wide array of sources including structured data files, Hive tables,
JSON, Parquet, external databases, HDFS, S3, etc. Additionally, through Spark
SQL’s external data sources API, DataFrames can be extended to support any thirdparty data formats or sources, including Avro, CSV, ElasticSearch, Cassandra, etc.
DataFrames are evaluated lazily, just like RDDs, while operations are automatically
parallelized and distributed on clusters. State-of-the-art optimization and code
generation is woven throughout the Spark SQL Catalyst optimizer utilizing a tree
transformation framework. DataFrames can be easily integrated with the rest of the
Hadoop ecosystem tools and frameworks via Spark Core and provides an API for
Python, Java, Scala, and R Programming (Fig. 6.17).
6.4.10 Creating a DataFrame
In order to start any Spark computation, a basic Spark session needs to be initialized
using the sparkR.session() command [33]. Code presented below is adapted
from the Spark http://spark.apache.org website.
From pyspark.sql import SparkSession
spark = SparkSession
.builder
.appName(“Python Spark example”)
.config(“myspark.config.option”, “myvalue”)
.getOrCreate()
Précédent

- 349/647

Suivant