14
F. Firouzi et al.
www.bestppt.com
Structured
CSV, Columnar Storage
(Parquet,
ORC). Strict data
model structure
Unstructured
Audio, video, images.
Meaningless without
adding some structure
SemiStructured
JSON, XML, sensor
data, social media,
device data, web logs.
Flexible data model
structure
Structured,
unstructured,
semi-structured
Terabytes of
data
Batch, realtime&
stream
Velocity
Volume
Variety
3 V’s
Before Big Data: Never visit data again
Big Data Idea: Store data & extract value from data
Fig. 1.2 The definition of big data
semi-structured data and the results of analyzing that data to gain insights. Doug
Laney defined big data as the three V’s (see Fig. 1.2):
• Volume – Storing large amounts of data.
• Velocity – The rate at which data is generated is high, so it must be stored or
processed quickly.
• Variety – There are many possible formats of the data, from structured numeric
to text, e-mails, video, audio, and so on.
Turning big data into tangible business insights is one of the major benefits of
IoT. Most well-known approaches for dealing with IoT data include:
• Analyzing data: Before data becomes useful in making decisions, it must be
analyzed. Traditional manual analytics, though powerful and informative, simply
will not be practical in the face of the staggering amount of data that IoT
will generate. Therefore, some automated analytics must be employed. These
analytics need to provide descriptive reports of the environment, visualizations,
dashboards, trigger alerts from data sources, and automated actions to be taken
based on the data. They will also be used to detect patterns in the data, predict
outcomes, and detect anomalies. There are open-source frameworks currently
available for performing automated analytics. The two main approaches are to
process the data in batches or to analyze the data as it is generated in real
time. Which technique to use depends on the context of the problem as well
as the resources available. The analytics can be run in a distributed fashion,
also called in the cloud or at the edge, in servers nearer the sensors. First,
the data is preprocessed, that is, duplicates are filtered out, and the data is
possibly reordered, aggregated, and most likely normalized. These and other
similar preprocessing tasks can be performed on the IoT device itself or on a
gateway device before it is sent upstream. The most common automated analytics
performed now are machine learning algorithms.
Précédent

- 23/647

Suivant