5 Machine Learning for IoT
275
Fig. 5.27 Entropy vs.
probability (target class is 1)
for a two-class variable (X)
1
0
1
0
P (X=1)
0.5
0.5
H(X)
of triangle class is 14/(16 + 14). As a result, the corresponding entropy can be
calculated as follows:
Entropy =
16
16 + 14
log 2
16
16 + 14
+
14
16 + 14
log 2
14
16 + 14
The entropy of a group that contains a single example class is zero. This means that
the group does not have any information. This also indicates that this group is not
a proper training set for the learning algorithm. In contrast, the entropy of a group
with 50% of either class is 1, which indicates that the group is a suitable training set
(i.e., the training set is balanced). Figure 5.27 illustrates the entropy vs. probability
of a two-class variable (e.g., red circle class and green triangle class). As you note,
the minimum of entropy is where the probability is equal to 0 or 1 (i.e., when we
have just red circles or green triangles). On the other hand, entropy rises to 1.0 at
a probability of 0.5 (maximum impurity) when the set is completely balanced (e.g.,
half of the examples are red circles and the other half examples are green triangles).
To be able to use entropy for feature selection, mutual information (MI) is
defined. MI shows how much information a variable has about another variable.
Larger mutual information (e.g., between the target (Y) and feature (X)) indicates
that the feature has more correlation with the target. Mutual information (MI) can
be calculated as
MI (Y, X) = H (Y ) + H (X) − H (Y, X)
In the above equation, H(Y, X) is a conditional entropy (entropy of a joint distribution):
H (Y, X) = −
i
p (y i , x i ) log 2 p (y i , x i )
Précédent

- 281/647

Suivant