6 Big Data
345
Reading data from a JSON file into a DataFrame is demonstrated in the code
below:
# Read the data file in JSON format
df = spark.read.json(“/user1/employees.json”)
# Displays the content of the DataFrame
df.show()
6.4.10.1 Example of Reading DataFrame from the Parquet File
dfParquet=spark.read.parquet
(“/user1/employees.parquet”)
display(dfParquet)
6.4.11 DataFrame Operations
DataFrames provide a domain-specific language for structured data manipulation in
Scala, Java, Python, and R. DataFrames are Dataset of Rows in the Scala and Java
APIs. These operations are also referred to as “untyped transformations” in contrast
to “typed transformations” typically associated with strongly typed Scala or Java
Datasets.
Basic examples of structured data processing using Datasets is demonstrated
below:
# Print the schema in a tree format
df.printSchema()
# Select only the “name” column
df.select(“name”).show()
# Select employees with salary greater than 3000
df.filter(df[’salary’] > 3000).show()
# Count people by salary
df.groupBy(“salary”).count().show()
The advantage of a SQL function on a SparkSession is that it enables applications
to run SQL queries programmatically and returns the result as a DataFrame [33].
Temporary views in Spark SQL are session-scoped and will disappear if the session
that creates it terminates. If there is a need for a temporary view to persist and
be shared among all sessions until the Spark application terminates, a global
temporary view should be created. A global temporary view is tied to a system
preserved database global_temp and must use the qualified name to refer it, e.g.,
SELECT * FROM global_temp.employee.
#Temporary view utilized to query the data
df.createOrReplaceTempView(“employees”)
sqlDF = spark.sql(“SELECT * FROM employees ”)
sqlDF.show()
345
Reading data from a JSON file into a DataFrame is demonstrated in the code
below:
# Read the data file in JSON format
df = spark.read.json(“/user1/employees.json”)
# Displays the content of the DataFrame
df.show()
6.4.10.1 Example of Reading DataFrame from the Parquet File
dfParquet=spark.read.parquet
(“/user1/employees.parquet”)
display(dfParquet)
6.4.11 DataFrame Operations
DataFrames provide a domain-specific language for structured data manipulation in
Scala, Java, Python, and R. DataFrames are Dataset of Rows in the Scala and Java
APIs. These operations are also referred to as “untyped transformations” in contrast
to “typed transformations” typically associated with strongly typed Scala or Java
Datasets.
Basic examples of structured data processing using Datasets is demonstrated
below:
# Print the schema in a tree format
df.printSchema()
# Select only the “name” column
df.select(“name”).show()
# Select employees with salary greater than 3000
df.filter(df[’salary’] > 3000).show()
# Count people by salary
df.groupBy(“salary”).count().show()
The advantage of a SQL function on a SparkSession is that it enables applications
to run SQL queries programmatically and returns the result as a DataFrame [33].
Temporary views in Spark SQL are session-scoped and will disappear if the session
that creates it terminates. If there is a need for a temporary view to persist and
be shared among all sessions until the Spark application terminates, a global
temporary view should be created. A global temporary view is tied to a system
preserved database global_temp and must use the qualified name to refer it, e.g.,
SELECT * FROM global_temp.employee.
#Temporary view utilized to query the data
df.createOrReplaceTempView(“employees”)
sqlDF = spark.sql(“SELECT * FROM employees ”)
sqlDF.show()
