4 Architecting IoT Cloud
187
robust compliance and security requirements. In contrast to Kafka and Flume, Nifi
is capable of handling data objects of differing sizes. NiFi comes with a userfriendly drag-and-drop, web-based user interface which allows users to visualize the
whole process and make necessary changes in real time. The core design concepts
underpinning Nifi are similar to the basic idea of Flow-Based Programming (FBP).
Below are the foundational Nifi concepts and components [9]:
• FlowFile (Information Packet) – It represents the data objects passing through
the system. For each FlowFile, Nifi tracks a set of key/value pair characteristics
of the object as well as its corresponding content (zero or more bytes).
• FlowFile Processor (Processor) – Processor handles a combination of data
transformation, routing, and system mediation. Processors are able to access
FlowFile characteristics as well its content. Processors are able to work on zero
or more FlowFiles in parallel. NiFi processors can either commit the result/out
or can roll back to their previous state (to address fault-tolerant issues). Nifi
comes with a wide assortment of processors (at the time of this writing more than
260 processors) including connectors for Kafka and Flume that may be dragged,
dropped, configured, and immediately put to work. There is also a possibility to
design and implement custom processors for Apache NiFi.
• Connection (Bounded Buffer) – Connections serve as a link between processors.
They function as queues and enable different processes to interact at various
rates. Queues can be arranged dynamically and may include upper bounds on the
load, allowing back pressure. Backpressure references scenarios where queues or
buffers are at capacity (full) and unable to receive new data. In such a case, the
backpressure mechanism ensures that new packets of data are not sent until the
data bottleneck has been resolved or the buffer is no longer full.
4.4.1.4 Elastic Logstash
Elastic Logstash is an open-source data ingestion framework that takes in data from
many sources concurrently, transforms the data, and sends the transformed data to
the Elasticsearch database (a NoSQL Database). Logstash is capable of ingesting
data of varying source, shape, and size from several sources such as logs, and
databases, in a constant, streaming manner. Logstash has a large variety of output
plugins, to support a range of different use cases. Additional Elasticsearch and
Logstash details will be discussed later in this chapter.
4.5 Data Processing Layer
In this section, we will discuss the modern data processing and big data architectures
created to manage huge amounts of data in order to extract the value of IoT data.
Before reviewing data processing architectures, some fundamental terms are defined
below [3, 5, 6, 10, 11]:
Précédent

- 194/647

Suivant