5 Machine Learning for IoT
261
method, the impact of the independent variable(s) on the dependent variable(s) is
analyzed. There are several well-known use cases for regression analysis:
• Predictive Analytics: Predictive analytics tries to model and predict future
behaviors by analyzing historical data. This technique has a wide range of
applications in IoT. As an example, predict maintenance predicts the time of
machine failures based on sensor readouts (e.g., vibration data) mounted in
machines in order to minimize downtime and to maximize productivity.
• Operation Efficiency: Regression analysis is used to optimize business processes
and assets (e.g., machine, workstation, laborer on the shop floor). For example,
on the production floor, IoT sensors mounted on machines can track inventory
consumption in real time. Regression analysis can forecast future behavior and
trigger automatic reordering or refill.
• Decision Support Systems: Regression helps make smarter and more accurate
decisions based on the available data.
• Error Correction: Regression analysis can be utilized to correct wrong decisions,
which are sometimes made based on some incorrect observations. For example,
a technical manager on a shop floor may believe that increasing the temperature
of a specific phase of the production line can increase the quality of the products.
However, readouts of IoT sensor and regression analysis indicate that this
assumption is not correct.
• New Insights: Regression analysis can reveal hidden patterns in data that are
difficult to uncover by conventional approaches. For example, regression analysis
techniques can be applied to IoT data of a production line to be able to find out
the relationship between environmental/operating conditions and the quality of
the products.
Formally, the regression task can be formulated as follows:
• There is a training set ((T = {(x (1) , y (1) ), . . . , (x (T) , y (T) )}) and we need to investigate the relationship (a mathematical equation) between the input (independent
variable(s) or features) x = (x 1 , . . . , x D ) and the output (dependent variable) y (i) .
• Note that in a training set of regression analysis, the labels y (i) are continuous.
This is one of the differences between regression tasks and classification tasks,
in which the y (i) are categorical.
Some important terminologies related to regression analysis are:
• Outliers: An outlier (Fig. 5.15) is a data point in a dataset in which a very high or a
very low value in comparison to other data points can be observed. An outlier can
be deemed to be an extreme value in the dataset. The presence of outliers leads
to less accurate results in regression; therefore, outliers are sometimes eliminated
through a preprocessing step.
• Multicollinearity: Multicollinearity is the situation of having a high level of correlation among the inputs. This means that the predictors (independent variables)
261
method, the impact of the independent variable(s) on the dependent variable(s) is
analyzed. There are several well-known use cases for regression analysis:
• Predictive Analytics: Predictive analytics tries to model and predict future
behaviors by analyzing historical data. This technique has a wide range of
applications in IoT. As an example, predict maintenance predicts the time of
machine failures based on sensor readouts (e.g., vibration data) mounted in
machines in order to minimize downtime and to maximize productivity.
• Operation Efficiency: Regression analysis is used to optimize business processes
and assets (e.g., machine, workstation, laborer on the shop floor). For example,
on the production floor, IoT sensors mounted on machines can track inventory
consumption in real time. Regression analysis can forecast future behavior and
trigger automatic reordering or refill.
• Decision Support Systems: Regression helps make smarter and more accurate
decisions based on the available data.
• Error Correction: Regression analysis can be utilized to correct wrong decisions,
which are sometimes made based on some incorrect observations. For example,
a technical manager on a shop floor may believe that increasing the temperature
of a specific phase of the production line can increase the quality of the products.
However, readouts of IoT sensor and regression analysis indicate that this
assumption is not correct.
• New Insights: Regression analysis can reveal hidden patterns in data that are
difficult to uncover by conventional approaches. For example, regression analysis
techniques can be applied to IoT data of a production line to be able to find out
the relationship between environmental/operating conditions and the quality of
the products.
Formally, the regression task can be formulated as follows:
• There is a training set ((T = {(x (1) , y (1) ), . . . , (x (T) , y (T) )}) and we need to investigate the relationship (a mathematical equation) between the input (independent
variable(s) or features) x = (x 1 , . . . , x D ) and the output (dependent variable) y (i) .
• Note that in a training set of regression analysis, the labels y (i) are continuous.
This is one of the differences between regression tasks and classification tasks,
in which the y (i) are categorical.
Some important terminologies related to regression analysis are:
• Outliers: An outlier (Fig. 5.15) is a data point in a dataset in which a very high or a
very low value in comparison to other data points can be observed. An outlier can
be deemed to be an extreme value in the dataset. The presence of outliers leads
to less accurate results in regression; therefore, outliers are sometimes eliminated
through a preprocessing step.
• Multicollinearity: Multicollinearity is the situation of having a high level of correlation among the inputs. This means that the predictors (independent variables)
