346
N. Balac
#DataFrame registered as a global temporary view
df.createGlobalTempView(“employees”)
spark.sql(“SELECT * FROM
global_temp.employees”).show()
#This will be available in the new Spark session
spark.newSession().sql(“SELECT * FROM
global_temp.employee”).show()
Remember to only use SELECT * in cases of small data sets similar to the ones
used in these illustrative examples; otherwise the WHERE clause should be utilized
to prevent the possibly very large amount of queried data.
Some additional examples of the queries enabled by Spark SQL are shown
below:
From pyspark.sql import functions as F
#Show all entries in the column named First name
df.select(“firstName”).show()
# Show all entries where salary >2000
df.select(df[’salary’] > 2000).show()
# Show first name and 0 or 1 depending if they are
# older or younger than 25
df.select(“firstName”, F.when(df.age > 25, 1)
.otherwise(0)).show()
6.4.12 Spark MLlib
Spark MLlib is a library containing various machine learning (ML) functionalities
optimized for the Spark computing framework. MLlib provides an extensive number
of machine learning algorithms and utilities including classification, regression,
clustering, association rules, sequential pattern mining, ensemble models, decomposition, topic modeling, and collaborative filtering [30]. In addition, MLlib supports
various functionalities such as feature extraction, model evaluation, and validation.
All of these methods are designed and optimized to scale across a Spark cluster.
Spark’s machine learning utilities enable construction of pipelines including tasks
that range from data ingest and feature transformations, data standardization,
normalization, summary statistics, dimensionality reduction, etc. to model building,
hyper-parameter tuning, and evaluation. Finally, Spark enables machine learning
persistence by saving and loading models and pipelines [37].
N. Balac
#DataFrame registered as a global temporary view
df.createGlobalTempView(“employees”)
spark.sql(“SELECT * FROM
global_temp.employees”).show()
#This will be available in the new Spark session
spark.newSession().sql(“SELECT * FROM
global_temp.employee”).show()
Remember to only use SELECT * in cases of small data sets similar to the ones
used in these illustrative examples; otherwise the WHERE clause should be utilized
to prevent the possibly very large amount of queried data.
Some additional examples of the queries enabled by Spark SQL are shown
below:
From pyspark.sql import functions as F
#Show all entries in the column named First name
df.select(“firstName”).show()
# Show all entries where salary >2000
df.select(df[’salary’] > 2000).show()
# Show first name and 0 or 1 depending if they are
# older or younger than 25
df.select(“firstName”, F.when(df.age > 25, 1)
.otherwise(0)).show()
6.4.12 Spark MLlib
Spark MLlib is a library containing various machine learning (ML) functionalities
optimized for the Spark computing framework. MLlib provides an extensive number
of machine learning algorithms and utilities including classification, regression,
clustering, association rules, sequential pattern mining, ensemble models, decomposition, topic modeling, and collaborative filtering [30]. In addition, MLlib supports
various functionalities such as feature extraction, model evaluation, and validation.
All of these methods are designed and optimized to scale across a Spark cluster.
Spark’s machine learning utilities enable construction of pipelines including tasks
that range from data ingest and feature transformations, data standardization,
normalization, summary statistics, dimensionality reduction, etc. to model building,
hyper-parameter tuning, and evaluation. Finally, Spark enables machine learning
persistence by saving and loading models and pipelines [37].
