5 Machine Learning for IoT
291
X=0
X=1
Out
X
Y
Z
1
1
0
1
1
1
0
0
1
0
1
0
E parent =1
E child1 = -(1/3)log(1/3) – (2/3)log(2.3) = 0.918
E child2 = 0
IG= 1- (3/4)(0.918)=0.311
Y=0
Y=1
E parent =1
E child1 = 0
E child2 = 0
IG= 1- (1/2)0 - (1/2)0 = 1
Fig. 5.46 A simple example of selecting the split attribute based on information gain (IG)
• Select a feature (attribute) for the root node of the tree and then create an edge
(i.e., a branch) for each value of the attribute.
• Split (divide) the data points (instances) into subsets (i.e., one for each branch
created in the previous step).
• Repeat the above two steps recursively for each branch (each subset).
• Stop the above recursion for a branch when all its instances are in the same class
(i.e., have the same label).
One important question that we need to address in the above algorithm is how to
select the root in each iteration. In other words, how can we identify which feature to
split upon? There are several methods to select the best attribute for splitting in each
step. In general, a good attribute is the one that splits the successor nodes as pure as
possible. It means that the corresponding branches contain mostly instances of one
class. This can be explained by entropy. The dividing procedure should decrease the
entropy because a good attribute (node) splits a set into subsets (regions) with the
higher homogeneity (higher entropy means there is a mix of different classes in each
region). To be able to implement this process, “information gain (IG)” is defined as
follows:
I nf ormation Gain = Entropy (P arent node)
− [Average Enropy (children nodes)]
Information gains represent the decrease in entropy after a split on an attribute. It
means that we need to select an attribute for splitting procedure that has the highest
information gain. In other words, an attribute with the highest information gain
leads to the most homogeneous branches (lower entropy or very pure subset). Let
us explain the abovementioned process with an example. Assume we have three
attributes (feature) and the target has two classes (see Fig. 5.46). If we split on X
attribute, the IG will be 0.31. On the other hand, if we select Y as our splitting
attribute, the IG will be 1. Therefore, we split based on Y.
291
X=0
X=1
Out
X
Y
Z
1
1
0
1
1
1
0
0
1
0
1
0
E parent =1
E child1 = -(1/3)log(1/3) – (2/3)log(2.3) = 0.918
E child2 = 0
IG= 1- (3/4)(0.918)=0.311
Y=0
Y=1
E parent =1
E child1 = 0
E child2 = 0
IG= 1- (1/2)0 - (1/2)0 = 1
Fig. 5.46 A simple example of selecting the split attribute based on information gain (IG)
• Select a feature (attribute) for the root node of the tree and then create an edge
(i.e., a branch) for each value of the attribute.
• Split (divide) the data points (instances) into subsets (i.e., one for each branch
created in the previous step).
• Repeat the above two steps recursively for each branch (each subset).
• Stop the above recursion for a branch when all its instances are in the same class
(i.e., have the same label).
One important question that we need to address in the above algorithm is how to
select the root in each iteration. In other words, how can we identify which feature to
split upon? There are several methods to select the best attribute for splitting in each
step. In general, a good attribute is the one that splits the successor nodes as pure as
possible. It means that the corresponding branches contain mostly instances of one
class. This can be explained by entropy. The dividing procedure should decrease the
entropy because a good attribute (node) splits a set into subsets (regions) with the
higher homogeneity (higher entropy means there is a mix of different classes in each
region). To be able to implement this process, “information gain (IG)” is defined as
follows:
I nf ormation Gain = Entropy (P arent node)
− [Average Enropy (children nodes)]
Information gains represent the decrease in entropy after a split on an attribute. It
means that we need to select an attribute for splitting procedure that has the highest
information gain. In other words, an attribute with the highest information gain
leads to the most homogeneous branches (lower entropy or very pure subset). Let
us explain the abovementioned process with an example. Assume we have three
attributes (feature) and the target has two classes (see Fig. 5.46). If we split on X
attribute, the IG will be 0.31. On the other hand, if we select Y as our splitting
attribute, the IG will be 1. Therefore, we split based on Y.
