338
N. Balac
Fig. 6.13 Spark’s approach to fast data sharing for queries
6.4.4 RDD Transformations and Actions
RDDs enable two main types of operations: transformations and actions. Transformation operations could be applied on RDDs and typically return another RDD,
while action operations trigger computation and return values.
Spark transformation is a function that produces new RDDs from existing
RDDs. It takes RDDs as input and produces one or more RDDs as output based
on the transformation applied. Each time a transformation is applied, a new RDD
is created. Note that the input RDDs cannot be changed due to the RDD design
requirement that they are immutable by nature.
The process of applying transformations builds a RDD lineage. It keeps track of
all of the parent RDDs of the final RDD(s). RDD lineage is also known as the RDD
operator graph or RDD dependency graph. It represents a logical execution plan
in the form of a Directed Acyclic Graph (DAG) of the entire set of parent RDDs.
RDD transformations are evaluated in a “lazy” manner, by performing the
computation only when an action requires a result to be returned. Therefore, they are
not executed immediately. Two of the most basic and often used transformations are
the map and filter. A map function iterates over every line in a RDD and applies that
function to every element of RDD, possibly enabling the flexibility that the input
and the return type of RDD may differ from each other. For example, the input
RDD type can be a string and after applying the map function the return RDD can
be Boolean. Filter functions return a new RDD, containing only the elements that
meet a predicate [32].
N. Balac
Fig. 6.13 Spark’s approach to fast data sharing for queries
6.4.4 RDD Transformations and Actions
RDDs enable two main types of operations: transformations and actions. Transformation operations could be applied on RDDs and typically return another RDD,
while action operations trigger computation and return values.
Spark transformation is a function that produces new RDDs from existing
RDDs. It takes RDDs as input and produces one or more RDDs as output based
on the transformation applied. Each time a transformation is applied, a new RDD
is created. Note that the input RDDs cannot be changed due to the RDD design
requirement that they are immutable by nature.
The process of applying transformations builds a RDD lineage. It keeps track of
all of the parent RDDs of the final RDD(s). RDD lineage is also known as the RDD
operator graph or RDD dependency graph. It represents a logical execution plan
in the form of a Directed Acyclic Graph (DAG) of the entire set of parent RDDs.
RDD transformations are evaluated in a “lazy” manner, by performing the
computation only when an action requires a result to be returned. Therefore, they are
not executed immediately. Two of the most basic and often used transformations are
the map and filter. A map function iterates over every line in a RDD and applies that
function to every element of RDD, possibly enabling the flexibility that the input
and the return type of RDD may differ from each other. For example, the input
RDD type can be a string and after applying the map function the return RDD can
be Boolean. Filter functions return a new RDD, containing only the elements that
meet a predicate [32].
