6 Big Data
339
After the transformation, the resultant RDD is different from its parent RDD. It
can be smaller if functions like count, filter, or sample are applied or larger when
functions like Cartesian or union are applied. Alternatively, it could remain the same
size when a map function is applied.
There are two types of transformations: narrow and wide. In narrow transformations, all the elements that are required to compute the records in single partition
reside in the single partition of the parent RDD. A limited subset of partitions is
used to calculate the result. Typical narrow transformations are considered to be
the result of map or filter functions. Alternatively, in a wide transformation, all the
elements that are required to compute the records in the single partition may reside
in numerous partitions of the parent RDD. Wide transformations typically result
from the application of join or intersection.
Transformations create RDDs from each other. However, in order to work with
the actual dataset, actions need to be performed. Actions are Spark RDD operations
that create non-RDD values. When an action is triggered and the result is calculated,
the new RDD is not automatically formed as it was the case with transformations.
The values of actions are stored to drivers or to the external storage system. In the
example below, the first line defines a base RDD from an external file.
#Create an RDD from a file on HDFS
text = sc.textFile(’hdfs://user1/mytext.txt’)
#Transform the RDD of lines into an RDD of words
mywords=text.flatMap(lambda line: line.split())
#Transform the RDD of words into an RDD of key/value pairs
mykeyvals=mywords.map(lambda word:(word,1))
RDD transformation vs. action
#Map Transformation example counting length
lineLength = text_map(Lambda s: len(s))
#Reduce Action Example
totalLength = lineLength.reduce (Lambda a, b: a+b)
#More RDD manipulation examples
# The saveAsTextFile action writes the contents of
# an RDD to the disk
rdd.saveAsTextFile(’hdfs://user1/myRDDoutput.txt’)
# The count action returns the number of elements
# in an RDD
numElements=rdd.count();
numElements;
print(numElements)
This approach enables executions to be optimized while operations are automatically parallelized and distributed on the clusters. These operations are able to handle
many machine learning algorithms that are iterative by nature, and the interactive ad
hoc queries needed for many analytics applications in an efficient manner. This is
Précédent

- 344/647

Suivant