192
F. Firouzi and B. Farahani
Fig. 4.11 Apache Storm
Spout:
Read subtitles
Bolt:
Separate words
Bolt:
Count names
Source file
Row text
Lines of text
Words
Count number of
unique words
mentioned in text
the form of a directed acyclic graph (DAG) that utilizes bolts and spouts. Storm’s
input stream is managed by a spout responsible for moving the data to a bolt that
then transforms the data. A bolt can move the data to a storage area or move it
to another bolt. Storm can be visualized as an interconnected chain of bolts that
somehow transform data collected by the spout (see Fig. 4.11) [14].
4.5.2.2 Apache Flink
Apache Flink is also an open-source streaming framework providing extensive realtime data processing pipeline capability. It is highly scalable and able to address
millions of events each second. Flink is designed based on the DataFlow model and
processes data as it arrives. One of the great features of Flink is its fault-tolerance
capability based on the checkpointing concept (i.e., saving internal states to external
sources/storage and recovering the state of the system in case of any failure).
Flink also provides an SQL API enabling individuals with a lack of programming
knowledge to develop a Flink solution much easier and faster.
4.5.2.3 Apache Spark
Apache Spark is an efficient, in-memory engine used for data processing. It
provides refined and articulate development APIs that enable data engineers and
Précédent

- 199/647

Suivant