5 Machine Learning for IoT
265
Underfitting
(high bias)
Right model
Overfitting
(high Variance)
Fig. 5.19 An example of underfitting (high bias) and overfitting (high variance) regression
Fig. 5.20 Relation of the
complexity of the model and
the error of the training and
test sets
Error
Optimum Complexity
Test Set &
Validation set
Training Set
Complexity of Model
predictions on the test dataset. This is also called the problem of high variance. On
the contrary, when our algorithm works so poorly that it is unable to fit even training
set well, it is said to be underfitting the data. It is also known as the problem of high
bias.
Take Fig. 5.19 as an example. In this diagram, the straight line corresponds
to a linear regression, which underfits the data and leads to large errors in the
training set. A regression model of the polynomial kernel (the middle subfigure
in Fig. 5.19) is the most suitable fit because it works well on both the training
and test datasets. Note that the regression model of higher-order polynomial kernel
(the right subfigure in Fig. 5.19) fits better on the training data, but it causes more
error in the test dataset (see Fig. 5.20). In this figure, by moving from the left
subfigure toward the right subfigure, the model tries to learn more details of the
input data. Although higher-order kernel (more complex regression models) leads
to higher accuracy/performance on training data, it may perform inaccurately on
unseen inputs (test data), which means it loses its generality and gets worse. In other
words, increasing the complexity of the model may decrease the training error, but
it may eventually increase the test error, as explained above (Fig. 5.20).
Regularization is a technique of adding information to the learning algorithm to
make the model more generalized. This, in turn, enhances the performance of the
model on the unseen data (test data) as well. By using regularization, the learning
algorithm is modified in a way that it acts more efficiently on unseen data. The
modified regression model contains another term/component in its cost function,
265
Underfitting
(high bias)
Right model
Overfitting
(high Variance)
Fig. 5.19 An example of underfitting (high bias) and overfitting (high variance) regression
Fig. 5.20 Relation of the
complexity of the model and
the error of the training and
test sets
Error
Optimum Complexity
Test Set &
Validation set
Training Set
Complexity of Model
predictions on the test dataset. This is also called the problem of high variance. On
the contrary, when our algorithm works so poorly that it is unable to fit even training
set well, it is said to be underfitting the data. It is also known as the problem of high
bias.
Take Fig. 5.19 as an example. In this diagram, the straight line corresponds
to a linear regression, which underfits the data and leads to large errors in the
training set. A regression model of the polynomial kernel (the middle subfigure
in Fig. 5.19) is the most suitable fit because it works well on both the training
and test datasets. Note that the regression model of higher-order polynomial kernel
(the right subfigure in Fig. 5.19) fits better on the training data, but it causes more
error in the test dataset (see Fig. 5.20). In this figure, by moving from the left
subfigure toward the right subfigure, the model tries to learn more details of the
input data. Although higher-order kernel (more complex regression models) leads
to higher accuracy/performance on training data, it may perform inaccurately on
unseen inputs (test data), which means it loses its generality and gets worse. In other
words, increasing the complexity of the model may decrease the training error, but
it may eventually increase the test error, as explained above (Fig. 5.20).
Regularization is a technique of adding information to the learning algorithm to
make the model more generalized. This, in turn, enhances the performance of the
model on the unseen data (test data) as well. By using regularization, the learning
algorithm is modified in a way that it acts more efficiently on unseen data. The
modified regression model contains another term/component in its cost function,
