184
F. Firouzi and B. Farahani
4.4.1 Data Ingestion Frameworks
4.4.1.1 Apache Flume
Apache Flume is a distributed system useful for capturing, aggregating, and routing
massive quantities of streaming data into a data store such as Hadoop. The details of
the Hadoop ecosystem will be discussed in Chap. 6. Flume includes several built-in
channels, sources, and sinks (data destinations). Flume also offers some features
necessary to be able to perform transformations (processing) of data during the
ingestion process. Flume can also be scaled horizontally in order to provide high
availability and scalability.
Apache Flume deployment includes starting one or multiple Flume agents. A
Flume Agent is a JVM process comprised of three elements: Flume Source, Flume
Channel, and Flume Sink (see Fig. 4.5) [7].
1. Flume Source – In Flume, an event is defined as a data unit consisting of a byte
payload and a set of optional string attributes. Data source is responsible for
ingesting the events triggered by an external source such as a Webserver or a
connected sensor.
2. Flume Channel – Afterward, the event is stored in one or more channels by Flume
Source. The stored data will stay in channel until it is consumed by Flume Sink.
3. Flume Sink – Finally, Flume Sink takes the event from channel storage and routes
it to an external depository (e.g., HDFS in Hadoop).
Note that agents are capable of being chained, so multiple Flume Agents can be
utilized. In such a case, Flume Sink of one agent sends the event on to the Flume
Source of the next Flume Agent in the chain.
4.4.1.2 Apache Kafka
Apache Kafka is an open source, high-throughput, distributed message
bus/queue/system to connect data consumers to data providers. In comparison
to Flume, Kafka provides greater scalability and more durable messaging. Kafka is
HDFS
(Hadoop)
Channel
Source
Sink
Client
Flume agent -JVM process
1
2
3
4
Fig. 4.5 Apache Flume Architecture
Précédent

- 191/647

Suivant