6 Big Data
347
6.4.13 MLlib Capabilities
MLlib’s capabilities enable utilization of the large number of major machine
learning algorithms including regression (linear, generalized linear, logistic), classification algorithms (including decision trees, random forest, gradient-boosted
tree, multilayer perceptron, support vector machine, naive Bayes, etc.), clustering
(K-means, K-medoids, bisecting k-means,) latent Dirichlet allocation, Gaussian
mixture model, and collaborative filtering. In addition, it supports feature extraction,
transformations, dimensionality reduction, selection, and the designing, constructing, evaluating, and tuning of machine learning pipelines.
There are many advantages of MLlib’s design including simplicity, scalability,
and compatibility. Spark’s APIs are simple by design and provide utilities that look
and feel like typical data science tools such as R and Python. Machine learning
methods can easily be executed with effective parameter tuning [38]. Additionally,
MLlib provides seamless scalability by enabling the execution of the ML methods
with minimal or no adjustment to the code on a large computing cluster. Spark is
compatible with R, Python pandas, scikit-learn, and many other prevalent ML tools.
Spark’s DataFrames and MLlib provide common data science tool integration with
existing workflows.
The goal of most machine learning experiments is to create an accurate model in
order to predict on future, unseen data. In order to accomplish this goal, a training
data set is used to “train” to fit the model, and a testing data set is used to evaluate
and validate the model obtained on the training data set.
Utilizing the PySpark MLlib features, traditional approaches to machine learning
can now be scaled to large and complex data sets. For example, we can use the
traditionally utilized Iris data set to demonstrate the capabilities of the MLlib to
develop predictive models on Spark.
from pyspark.ml import Pipeline
from pyspark.ml.classification import
DecisionTreeClassifier
from pyspark.ml.evaluation import
MulticlassClassificationEvaluator
from pyspark.ml.feature import
StringIndexer, VectorIndexer
from pyspark.ml.linalg import Vectors
from pyspark.ml.feature import VectorAssembler
#Get Spark Context from Spark Session
SpContext = SpSession.sparkContext
#Load the Iris.CSV file
df = spark.read.csv(“Iris.data”, inferSchema=True)
.toDF(“sepLenght”, “sepWidth”,
“petLenght”, “petWidth”, “class”)
#Print the first 10 rows of the DataFrame
df.show(10)
347
6.4.13 MLlib Capabilities
MLlib’s capabilities enable utilization of the large number of major machine
learning algorithms including regression (linear, generalized linear, logistic), classification algorithms (including decision trees, random forest, gradient-boosted
tree, multilayer perceptron, support vector machine, naive Bayes, etc.), clustering
(K-means, K-medoids, bisecting k-means,) latent Dirichlet allocation, Gaussian
mixture model, and collaborative filtering. In addition, it supports feature extraction,
transformations, dimensionality reduction, selection, and the designing, constructing, evaluating, and tuning of machine learning pipelines.
There are many advantages of MLlib’s design including simplicity, scalability,
and compatibility. Spark’s APIs are simple by design and provide utilities that look
and feel like typical data science tools such as R and Python. Machine learning
methods can easily be executed with effective parameter tuning [38]. Additionally,
MLlib provides seamless scalability by enabling the execution of the ML methods
with minimal or no adjustment to the code on a large computing cluster. Spark is
compatible with R, Python pandas, scikit-learn, and many other prevalent ML tools.
Spark’s DataFrames and MLlib provide common data science tool integration with
existing workflows.
The goal of most machine learning experiments is to create an accurate model in
order to predict on future, unseen data. In order to accomplish this goal, a training
data set is used to “train” to fit the model, and a testing data set is used to evaluate
and validate the model obtained on the training data set.
Utilizing the PySpark MLlib features, traditional approaches to machine learning
can now be scaled to large and complex data sets. For example, we can use the
traditionally utilized Iris data set to demonstrate the capabilities of the MLlib to
develop predictive models on Spark.
from pyspark.ml import Pipeline
from pyspark.ml.classification import
DecisionTreeClassifier
from pyspark.ml.evaluation import
MulticlassClassificationEvaluator
from pyspark.ml.feature import
StringIndexer, VectorIndexer
from pyspark.ml.linalg import Vectors
from pyspark.ml.feature import VectorAssembler
#Get Spark Context from Spark Session
SpContext = SpSession.sparkContext
#Load the Iris.CSV file
df = spark.read.csv(“Iris.data”, inferSchema=True)
.toDF(“sepLenght”, “sepWidth”,
“petLenght”, “petWidth”, “class”)
#Print the first 10 rows of the DataFrame
df.show(10)
