5 Machine Learning for IoT
269
Fig. 5.23 Comparison of
geometric interpretations of
lasso (L1), ridge (L2), and
elastic net
Ridge
Lasso
Elastic Net
Elastic net is particularly useful when we have several correlated features. When two
features have correlation, lasso selects one of them randomly, whereas the elastic
net considers both. Therefore, it is also more stable in many cases. For example,
imagine a problem in which there are two correlated features. In this case, lasso
chooses one of them randomly, but the elastic net takes both into account. Like
the ridge regression, the elastic net algorithm is more stable in comparison to other
methods for most of the real-world problems. Figure 5.23 depicts the comparison of
these algorithms [6, 7].
5.2.3 Bayesian Linear Regression
In linear regression, we have just one output value (y) for a given input. However,
Bayesian has another point of view. In this approach, y is not a single value, but it
is taken from a probability distribution. Recall that the linear regression approach
models the relation between input data (features) and the output (target) by the
following equation:
y = β
T X + ε
in which the response is produced by multiplying model parameters (i.e., weights:
β) by the input (i.e., X) plus the model error (ε), which might be caused by random
sampling noise or latent variables. In the ordinary least squares (OLS) approach, the
model parameter (weights) can be determined by minimizing the sum of squared
errors (Eq. 5.1). However, Bayesian linear regression uses a statistical approach
based on probability distribution such as Gaussian distribution to model the mapping
function between inputs (features) and output:
y ∼ N
β
T X, σ
2
269
Fig. 5.23 Comparison of
geometric interpretations of
lasso (L1), ridge (L2), and
elastic net
Ridge
Lasso
Elastic Net
Elastic net is particularly useful when we have several correlated features. When two
features have correlation, lasso selects one of them randomly, whereas the elastic
net considers both. Therefore, it is also more stable in many cases. For example,
imagine a problem in which there are two correlated features. In this case, lasso
chooses one of them randomly, but the elastic net takes both into account. Like
the ridge regression, the elastic net algorithm is more stable in comparison to other
methods for most of the real-world problems. Figure 5.23 depicts the comparison of
these algorithms [6, 7].
5.2.3 Bayesian Linear Regression
In linear regression, we have just one output value (y) for a given input. However,
Bayesian has another point of view. In this approach, y is not a single value, but it
is taken from a probability distribution. Recall that the linear regression approach
models the relation between input data (features) and the output (target) by the
following equation:
y = β
T X + ε
in which the response is produced by multiplying model parameters (i.e., weights:
β) by the input (i.e., X) plus the model error (ε), which might be caused by random
sampling noise or latent variables. In the ordinary least squares (OLS) approach, the
model parameter (weights) can be determined by minimizing the sum of squared
errors (Eq. 5.1). However, Bayesian linear regression uses a statistical approach
based on probability distribution such as Gaussian distribution to model the mapping
function between inputs (features) and output:
y ∼ N
β
T X, σ
2
