Relative Entropy
31
conditional probability instead of the original probability. 17 This is a good problem
in the sense that it teaches us.
Relative Entropy
We introduced relative entropy D KL in this chapter. This quantity is one of the
most important functions in the field of machine learning. Here, we will show the
properties quoted in this chapter: for a probability distribution p(x), q(x), we have
D KL (p||q) ≥ 0, and only when D KL (p||q) = 0 the probability distributions match,
p(x) = q(x). In addition, as an exercise, we will explain a calculation example
using the Gaussian distribution and some unexpected geometric structure behind it.
Definition and properties
For probability distributions p(x) and q(x), we define
D KL (p||q) :=
dx p(x) log
p(x)
q(x)
(2.42)
and call it relative entropy. The relative entropy satisfies
D KL (p||q) ≥ 0, D KL (p||q) = 0 ⇔ ∀x, p(x) = q(x).
(2.43)
There are several ways to check the above properties, and in this column we prove
them in the following way. First, write the relative entropy as
D KL (p||q) =
dx
p(x) log
p(x)
q(x)
+ q(x) − p(x)
.
(2.44)
Here we used
dx q(x) =
dx p(x) = 1. We simply added unity and subtracted
unity. And we can transform it as
(2.44) =
dx p(x)
q(x)
p(x)
− 1
− log
q(x)
p(x)
.
(2.45)
Using (see Fig. 2.3)
(X − 1) ≥ log X ,
(2.46)
17 For example, even if you have a favorite restaurant in your neighborhood, if you hear rumors that
the food at another restaurant is really delicious, then you will be tempted to go there. Considering
the conditional probability is something close to this simple feeling.
31
conditional probability instead of the original probability. 17 This is a good problem
in the sense that it teaches us.
Relative Entropy
We introduced relative entropy D KL in this chapter. This quantity is one of the
most important functions in the field of machine learning. Here, we will show the
properties quoted in this chapter: for a probability distribution p(x), q(x), we have
D KL (p||q) ≥ 0, and only when D KL (p||q) = 0 the probability distributions match,
p(x) = q(x). In addition, as an exercise, we will explain a calculation example
using the Gaussian distribution and some unexpected geometric structure behind it.
Definition and properties
For probability distributions p(x) and q(x), we define
D KL (p||q) :=
dx p(x) log
p(x)
q(x)
(2.42)
and call it relative entropy. The relative entropy satisfies
D KL (p||q) ≥ 0, D KL (p||q) = 0 ⇔ ∀x, p(x) = q(x).
(2.43)
There are several ways to check the above properties, and in this column we prove
them in the following way. First, write the relative entropy as
D KL (p||q) =
dx
p(x) log
p(x)
q(x)
+ q(x) − p(x)
.
(2.44)
Here we used
dx q(x) =
dx p(x) = 1. We simply added unity and subtracted
unity. And we can transform it as
(2.44) =
dx p(x)
q(x)
p(x)
− 1
− log
q(x)
p(x)
.
(2.45)
Using (see Fig. 2.3)
(X − 1) ≥ log X ,
(2.46)
17 For example, even if you have a favorite restaurant in your neighborhood, if you hear rumors that
the food at another restaurant is really delicious, then you will be tempted to go there. Considering
the conditional probability is something close to this simple feeling.
