1 IoT Fundamentals: Definitions, Architectures, Challenges, and Promises
15
• Machine learning (ML): Traditional mathematical statistical models analyze data
by fitting the data to a model. Then the model is used to make predictions. This
is a difficult process to follow especially when the data is dynamic or has many
variables or the important points of the data are unknown. Machine learning is
an algorithmic approach where the important parameters are extracted from the
data in a process called learning. The data itself provides the structure of the
mathematical model. Machine learning techniques can be applied to historical
data or data taken in real time. The main way to think of it is that machine
learning finds patterns or relationships or key variables in data. The model that is
learned can be updated over time as more data is collected. One of the important
applications of machine learning in the context of IoT is about finding patterns
in the data, so that anomalies can be found quickly. Traditionally, anomalies
were detected when certain values crossed thresholds. Machine learning allows
for more complex patterns in the data to be identified as anomalous, therefore
increasing speed and accuracy in detecting problems. The machine-learningdriven intelligence in IoT can be used for predictive analytics (what will happen),
prescriptive analytics (what should we do next), and adaptive or continuous
analytics (how can we adapt to the latest changes).
• Edge analytics: When analytics is applied at the edge of the network close to the
IoT devices that generate the input data, it is referred to as edge analytics. Since
network traffic is reduced, this is an attractive approach to reduce bandwidth and
the latency from data gathering to a useful result. One drawback is that more
processing power is needed in the devices and close to them, and cost or the
particulars of the environment or device may make this prohibitive. On the other
hand, sending large amounts of data across a network into the cloud may also be
too expensive. Often a hybrid of edge and upstream analytics in the cloud is used
to mitigate these costs.
• Real-time analytics: Any time that data is collected and immediately analyzed is
known as real-time analytics. This is the best choice when a delay in the results of
the analysis would reduce the value of the data. Time series data, rolling metrics,
running averages, and any other occasion where the window of time analysis
needs to be controlled are also good candidates for real-time analytics. Some
real-time analytics frameworks available include Apache Storm, Apache Spark,
and Flink frameworks.
• Distributed analytics: When the data sets are particularly large, too large to be
handled by a single node (server), then distributed analytics can be used. As
the name implies, the analysis tasks can be broken up and spread out to several
compute nodes, possibly across multiple databases. If the data allows, it could
be bucketed by time period and thereby effectively split up in order to make it
more manageable. This is also an example of batch processing. Hadoop provides
an ecosystem of frameworks for performing analytics. Apache Hadoop is used
for batch processing and uses the MapReduce engine to process distributed
data. Hadoop is a good open-source framework and one of the first to become
available. It is used successfully for historical data analytics.
Précédent

- 24/647

Suivant