348
N. Balac
Index the label by converting class to numeric using StringIndexer:
class_index = StringIndexer(inputCol=“class”,
outputCol=“classIndex”)
df = class_index.fit(df).transform(df)
# Split the data into train and test
(trainingData, testData) =
df.randomSplit([0.8, 0.2])
#train the Decision Tree Model
dt = DecisionTreeClassifier(labelCol=“ClassIndex”,
featuresCol=“features”)
model = dt.fit(trainingData)
predictions = model.transform(testData)
#Evaluate models’ accuracy
evaluation = MulticlassClassificationEvaluator(
labelCol=“labelIndex”, predictionCol=“prediction”,
metricName=“accuracy”)
accuracy = evaluation.evaluate(predictions)
#Print models’ error
print(“Test Error = %g ” % (1.0 - accuracy))
#Print Model summary
Print(model)
Spark enables solving multiple data problems on one platform, from analytics to
graph analysis and machine learning. The Spark ecosystem also provides a utility for
graph computations called GraphX in addition to streaming and real-time interactive
query processing with Spark SQL and DataFrames [36].
6.4.14 Spark Streaming
Spark Streaming is a Spark component that enables processing of live streams of
data and enables scalable, high-throughput, fault-tolerant data stream processing.
6.4.15 Intro to Batch and Stream Processing
Before looking into how specifics of how Spark Streaming works, the difference
between batch and stream processing should be defined. Typically, batch processing
collects a large volume of data elements into a group at once. The entire group is
then processed simultaneously in a batch at a specified time. The time of batch
computation can be quantified in a number of ways. The computation time can
be determined on a prespecified scheduled time interval or on specific triggered
condition including a number of elements of or amount of data collected. Batch
N. Balac
Index the label by converting class to numeric using StringIndexer:
class_index = StringIndexer(inputCol=“class”,
outputCol=“classIndex”)
df = class_index.fit(df).transform(df)
# Split the data into train and test
(trainingData, testData) =
df.randomSplit([0.8, 0.2])
#train the Decision Tree Model
dt = DecisionTreeClassifier(labelCol=“ClassIndex”,
featuresCol=“features”)
model = dt.fit(trainingData)
predictions = model.transform(testData)
#Evaluate models’ accuracy
evaluation = MulticlassClassificationEvaluator(
labelCol=“labelIndex”, predictionCol=“prediction”,
metricName=“accuracy”)
accuracy = evaluation.evaluate(predictions)
#Print models’ error
print(“Test Error = %g ” % (1.0 - accuracy))
#Print Model summary
Print(model)
Spark enables solving multiple data problems on one platform, from analytics to
graph analysis and machine learning. The Spark ecosystem also provides a utility for
graph computations called GraphX in addition to streaming and real-time interactive
query processing with Spark SQL and DataFrames [36].
6.4.14 Spark Streaming
Spark Streaming is a Spark component that enables processing of live streams of
data and enables scalable, high-throughput, fault-tolerant data stream processing.
6.4.15 Intro to Batch and Stream Processing
Before looking into how specifics of how Spark Streaming works, the difference
between batch and stream processing should be defined. Typically, batch processing
collects a large volume of data elements into a group at once. The entire group is
then processed simultaneously in a batch at a specified time. The time of batch
computation can be quantified in a number of ways. The computation time can
be determined on a prespecified scheduled time interval or on specific triggered
condition including a number of elements of or amount of data collected. Batch
