5 Machine Learning for IoT
295
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
+
+
-
+
-
+ +
+
-
-
-
+
+
+
-
+ +
+
+
-
-
-
-
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
+
+ +
-
+
+
+
-
-
-
+
-
+ +
+
+
-
-
-
Original Dataset | D1
Update weights| D2
Update weights| D3
Trained model| M1
Trained model| M2
Trained model| M3
Final model = f (M1, M2, M3)
Fig. 5.49 The overall procedure of the AdaBoost algorithm
f (x) =
T
t=1
α t h t (x)
in which the h t (x) are the basis classifiers. Here, the weak classifiers are single split
decision trees (called decision stumps). In the beginning, all training data have equal
weight. AdaBoost modifies the weights in a way that difficult to classify instances
get more weight and adds new weak classifiers sequentially with the focus on more
difficult instances. In this regard, each classifier (weak classifier) is trained by taking
a random subset of the training set, and then, AdaBoost assigns higher weights to
misclassified training items. As a result, this misclassified item will have a higher
probability to appear in the next training subset for the next classifier.
A simple example of AdaBoost procedure is depicted in Fig. 5.49. This figure
explains how AdaBoost updates the weights in each step to construct a final model
by a linear combination of several weak classifiers. Bigger weights are illustrated
by larger signs and smaller weights are shown by smaller signs.
Gradient Boosting
There is another type of boosting that is contrary to AdaBoost and works on the basis
of training on the remaining errors (or pseudo-residuals) of stronger classifiers. This
approach is known as gradient boosting, in which at each training iteration, a weak
classifier is fitted on the computed pseudo-residuals. Then the effect of this weak
classifier on the performance of the stronger one is computed based on a gradient
descent optimization process.
295
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
+
+
-
+
-
+ +
+
-
-
-
+
+
+
-
+ +
+
+
-
-
-
-
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
-
+
-
+ +
+
+
-
-
-
+
+ +
-
+
+
+
-
-
-
+
-
+ +
+
+
-
-
-
Original Dataset | D1
Update weights| D2
Update weights| D3
Trained model| M1
Trained model| M2
Trained model| M3
Final model = f (M1, M2, M3)
Fig. 5.49 The overall procedure of the AdaBoost algorithm
f (x) =
T
t=1
α t h t (x)
in which the h t (x) are the basis classifiers. Here, the weak classifiers are single split
decision trees (called decision stumps). In the beginning, all training data have equal
weight. AdaBoost modifies the weights in a way that difficult to classify instances
get more weight and adds new weak classifiers sequentially with the focus on more
difficult instances. In this regard, each classifier (weak classifier) is trained by taking
a random subset of the training set, and then, AdaBoost assigns higher weights to
misclassified training items. As a result, this misclassified item will have a higher
probability to appear in the next training subset for the next classifier.
A simple example of AdaBoost procedure is depicted in Fig. 5.49. This figure
explains how AdaBoost updates the weights in each step to construct a final model
by a linear combination of several weak classifiers. Bigger weights are illustrated
by larger signs and smaller weights are shown by smaller signs.
Gradient Boosting
There is another type of boosting that is contrary to AdaBoost and works on the basis
of training on the remaining errors (or pseudo-residuals) of stronger classifiers. This
approach is known as gradient boosting, in which at each training iteration, a weak
classifier is fitted on the computed pseudo-residuals. Then the effect of this weak
classifier on the performance of the stronger one is computed based on a gradient
descent optimization process.
