22
2 Introduction to Machine Learning
take the following strategy:
I. Define the probability Q J (x, d) depending on the parameter J
II. Adjust J to make Q J (x, d) as close as possible to P (x, d)
Stage II is called training or learning.
Suppose the model Q J (x, d) is defined in some way. In order to bring this closer to
P (x, d) at stage II, we need a function that measures the difference between them.
A suitable one is relative entropy: 6
D KL (P ||Q J ) =
x,d
P (x, d) log
P (x, d)
Q J (x, d)
.
(2.8)
Many readers may be unfamiliar with the || notation used here. This is just a notation
to emphasize that the arguments P , Q J are asymmetric. The relative entropy is nonnegative, i.e., D KL ≥ 0, where the equality holds if and only if P (x, d) = Q J (x, d)
for any (x, d) (see the column of this chapter). Therefore, if we can adjust (learn)
J and reduce D KL , learning will proceed. Equation (2.8) is called generalization
error. 7 Generalization is a technical term used to describe the state in which a
machine can adapt to unknown situations, as described at the beginning of this
chapter.
Importance of the number of data
Actual learning has various difficulties. First of all, it is impossible to calculate the
value or gradient of (2.8) because we never know the specific expression of P (x, d).
In realistic situations, we use the approximate probability ˆ
P (x, d) derived from the
data (this is called empirical probability) such as (2.4), and consider
D KL ( ˆ
P ||Q J )
(2.9)
which is called empirical error. 8 From this, it can be intuitively understood that
reliable learning results cannot be obtained unless the number of data is large
enough. For example, (2.4), the calculation was performed with a set of 1000 data.
6 As mentioned in Chap. 1, relative entropy is also called Kullback–Leibler divergence. Although
it measures the “distance,” it does not satisfy the axiom of symmetry for the distance, so it is called
divergence.
7 In general, “generalization error” often refers to the expectation value of the error function (which
we will describe later). As shown later, they are essentially the same thing.
8 This is the same as using maximum likelihood estimation, just as we used it when we introduced
relative entropy in Chap. 1.
2 Introduction to Machine Learning
take the following strategy:
I. Define the probability Q J (x, d) depending on the parameter J
II. Adjust J to make Q J (x, d) as close as possible to P (x, d)
Stage II is called training or learning.
Suppose the model Q J (x, d) is defined in some way. In order to bring this closer to
P (x, d) at stage II, we need a function that measures the difference between them.
A suitable one is relative entropy: 6
D KL (P ||Q J ) =
x,d
P (x, d) log
P (x, d)
Q J (x, d)
.
(2.8)
Many readers may be unfamiliar with the || notation used here. This is just a notation
to emphasize that the arguments P , Q J are asymmetric. The relative entropy is nonnegative, i.e., D KL ≥ 0, where the equality holds if and only if P (x, d) = Q J (x, d)
for any (x, d) (see the column of this chapter). Therefore, if we can adjust (learn)
J and reduce D KL , learning will proceed. Equation (2.8) is called generalization
error. 7 Generalization is a technical term used to describe the state in which a
machine can adapt to unknown situations, as described at the beginning of this
chapter.
Importance of the number of data
Actual learning has various difficulties. First of all, it is impossible to calculate the
value or gradient of (2.8) because we never know the specific expression of P (x, d).
In realistic situations, we use the approximate probability ˆ
P (x, d) derived from the
data (this is called empirical probability) such as (2.4), and consider
D KL ( ˆ
P ||Q J )
(2.9)
which is called empirical error. 8 From this, it can be intuitively understood that
reliable learning results cannot be obtained unless the number of data is large
enough. For example, (2.4), the calculation was performed with a set of 1000 data.
6 As mentioned in Chap. 1, relative entropy is also called Kullback–Leibler divergence. Although
it measures the “distance,” it does not satisfy the axiom of symmetry for the distance, so it is called
divergence.
7 In general, “generalization error” often refers to the expectation value of the error function (which
we will describe later). As shown later, they are essentially the same thing.
8 This is the same as using maximum likelihood estimation, just as we used it when we introduced
relative entropy in Chap. 1.
