292
F. Firouzi et al.
The other challenge in the construction of decision trees is the stop criteria of
splitting procedure, in other words, when we need to stop the splitting. One possible
approach is to continue the splitting until each leaf node in the decision tree is
completely pure (i.e., only the instances/examples of one class are in leaf). This
approach might not be very efficient for all the applications, because it creates a
large number of small regions in the feature space. From a technical perspective,
this increases the overfitting. Pruning techniques are a great solution to tackle this
issue.
• Pre-pruning: We stop growing the tree when the performance of the classifier is
higher than a predefined threshold.
• Post-Pruning: First, we grow the complete decision tree. Although a complete
tree can classify all the training data points correctly, it may suffer from
overfitting. Next, a pruning algorithm is iteratively applied to simplify the tree
by removing some of its nodes. This iterative algorithm decides to remove some
of the nodes if the increase in entropy is below a predefined threshold.
5.4.7 Ensembles
An ensemble machine learning combines several weak models in order to create one
single strong meta-model to be able to tack high bias (underfitting) and high variance
(overfitting) issues. There are two broad categories of ensemble techniques, namely,
bootstrap aggregating (bagging) and boosting.
5.4.7.1 Bootstrap Aggregating (Bagging)
Bootstrap aggregating, which is also called bagging, is an effective technique to
reduce the model overfitting and to handle unstable datasets. It also improves the
performance of training on a dataset with a limited number of training data. For
example, bagging can be used to combine multiple decision trees as a forest model
to obtain a stronger classifier.
This method generates n small training set from the original dataset. Each of
them is called one sample. These samples are produced by random sampling of the
input dataset with replacement, as illustrated in Fig. 5.47. Let us explain the concept
of sampling with replacement with a simple example. Suppose we have three names
and we need to sample two. The names are Farshad, Sani, and Krish. We put these
three names in a hat and randomly choose one of them. Then we put the name back
into the hat and select another name. The possibilities of two-sample names are
(Farshad, Farshad), (Farshad, Sani), (Sani, Sani), and so on.
When all samples are constructed by random sampling with replacement technique, we independently learn and build n ensembles (i.e., individual learning
models) corresponding to n samples. Each of these weak models is called the
F. Firouzi et al.
The other challenge in the construction of decision trees is the stop criteria of
splitting procedure, in other words, when we need to stop the splitting. One possible
approach is to continue the splitting until each leaf node in the decision tree is
completely pure (i.e., only the instances/examples of one class are in leaf). This
approach might not be very efficient for all the applications, because it creates a
large number of small regions in the feature space. From a technical perspective,
this increases the overfitting. Pruning techniques are a great solution to tackle this
issue.
• Pre-pruning: We stop growing the tree when the performance of the classifier is
higher than a predefined threshold.
• Post-Pruning: First, we grow the complete decision tree. Although a complete
tree can classify all the training data points correctly, it may suffer from
overfitting. Next, a pruning algorithm is iteratively applied to simplify the tree
by removing some of its nodes. This iterative algorithm decides to remove some
of the nodes if the increase in entropy is below a predefined threshold.
5.4.7 Ensembles
An ensemble machine learning combines several weak models in order to create one
single strong meta-model to be able to tack high bias (underfitting) and high variance
(overfitting) issues. There are two broad categories of ensemble techniques, namely,
bootstrap aggregating (bagging) and boosting.
5.4.7.1 Bootstrap Aggregating (Bagging)
Bootstrap aggregating, which is also called bagging, is an effective technique to
reduce the model overfitting and to handle unstable datasets. It also improves the
performance of training on a dataset with a limited number of training data. For
example, bagging can be used to combine multiple decision trees as a forest model
to obtain a stronger classifier.
This method generates n small training set from the original dataset. Each of
them is called one sample. These samples are produced by random sampling of the
input dataset with replacement, as illustrated in Fig. 5.47. Let us explain the concept
of sampling with replacement with a simple example. Suppose we have three names
and we need to sample two. The names are Farshad, Sani, and Krish. We put these
three names in a hat and randomly choose one of them. Then we put the name back
into the hat and select another name. The possibilities of two-sample names are
(Farshad, Farshad), (Farshad, Sani), (Sani, Sani), and so on.
When all samples are constructed by random sampling with replacement technique, we independently learn and build n ensembles (i.e., individual learning
models) corresponding to n samples. Each of these weak models is called the
