122
6 Unsupervised Deep Learning
What quantity should be considered using these two? First, suppose the generator
draws some image:
x fake ∼ Q G ∗ (x).
(6.99)
Using this image, what is the probability Q J ∗ (d|x fake ) of the classification label d?
For simplicity, let d = (dog, cat). As a first example, suppose the generated image
is completely useless, and x fake looks like a mysterious animal between dogs and
cats. Then the image looks like a cat and like a dog, so
Q J ∗ (d = dog|x fake ) ≈
1
2
, Q J ∗ (d = cat|x fake ) ≈
1
2
.
(6.100)
On the other hand, if x fake is an image which looks really like a dog, we should have
Q J ∗ (d = dog|x fake ) ≈ 1, Q J ∗ (d = cat|x fake ) ≈ 0 .
(6.101)
The entropy in each case is
S(x fake ) = −
d
Q J ∗ (d|x fake ) log Q J ∗ (d|x fake )
≈
log 2 (x fake is a bad image (6.100)),
0
(x fake is a good image (6.101)).
(6.102)
Namely, the closer S(x fake ) is to 0, the more reality it has. So, this quantity is the
appropriate one to calculate.
But here is a trap. Equation (6.102) certainly measures the reality, but it only
measures goodness for a single x fake . For example, by sampling from the generator,
we obtain many images
x fake1 , x fake2 , . . . , ∼ Q G ∗ (x) ,
(6.103)
and if most of the images are similar dog images, the individual (6.102) values
are certainly large, but it is not possible to create a new image, so, resultantly, the
machine has not generalized. In other words, generators are required to generate
realistic images that are “sufficiently diverse.” Entropy can also be used to measure
such diversity. We just need to calculate the entropy of the expectation value of the
classification probability when various x fake are input,
Q(d) =
dxQ J ∗ (d|x)Q G ∗ (x) = =Q J ∗ (d|x) x∼Q G ∗ (x) .
(6.104)
Précédent

- 130/211

Suivant