6.4 Generalization in Generative Models
123
Explaining again with the example of dogs and cats, if Q G ∗ (x) generates both dog
and cat images equally, the probability for x fake to be classified as dog or cat is
1
2 ,
Q(d = dog) ≈
1
2
, Q(d = cat) . ≈
1
2
.
(6.105)
On the other hand, if Q G ∗ (x) generates only dogs as in the case above, we have
Q(d = dog) ≈ 1, Q(d = cat) ≈ 0 .
(6.106)
Then the entropy of Q(d) is
S = −
d
Q(d) log Q(d)
≈
log 2 (Q G ∗ is diverse, (6.105)),
0
(Q G ∗ is not diverse, (6.106)).
(6.107)
We see that the larger S, the more diverse.
Taken together, a good generator should decrease (6.102) and increase (6.107).
Since entropy is additive, the “score” of the generator can be thought of as the
difference. Then this can be written as an expectation value of a relative entropy. 19
S − −S(x) x∼Q G ∗ (x)
=
d
J ∗ (d|x) log Q J ∗ (d|x) x∼Q G ∗ (x)
Transform this into an integration
− Q(d)
(6.104)
log Q(d)
=
d
dx Q G ∗ (x)Q J ∗ (d|x) log Q J ∗ (d|x) −
dx Q G ∗ (x)Q J ∗ (d|x) log Q(d)
=
dx Q G ∗ (x)
d
Q J ∗ (d|x) log
Q J ∗ (d|x)
Q(d)
=
dx Q G ∗ (x)D KL
Q J ∗ (d|x)
Q(d)
=
D KL
Q J ∗ (d|x)
Q(d)
x∼Q G ∗ (x)
.
(6.109)
19 Yet another transform
(6.109) =
dx
d
Q J ∗ G ∗ (x, d) log
Q J ∗ G ∗ (x, d)
Q G ∗ (x)Q(d)
, Q J ∗ G ∗ (x, d) = Q J ∗ (d|x)Q G ∗ (x)
(6.108)
provides a quantity called mutual information between the generated image and the classification
labels. Reference [93] used this to hack IS, that is, generate images that have an unusually large
value of the IS (although images that make little sense to the human eyes). As seen from this
example, it is not always true that the higher the IS, the better.
123
Explaining again with the example of dogs and cats, if Q G ∗ (x) generates both dog
and cat images equally, the probability for x fake to be classified as dog or cat is
1
2 ,
Q(d = dog) ≈
1
2
, Q(d = cat) . ≈
1
2
.
(6.105)
On the other hand, if Q G ∗ (x) generates only dogs as in the case above, we have
Q(d = dog) ≈ 1, Q(d = cat) ≈ 0 .
(6.106)
Then the entropy of Q(d) is
S = −
d
Q(d) log Q(d)
≈
log 2 (Q G ∗ is diverse, (6.105)),
0
(Q G ∗ is not diverse, (6.106)).
(6.107)
We see that the larger S, the more diverse.
Taken together, a good generator should decrease (6.102) and increase (6.107).
Since entropy is additive, the “score” of the generator can be thought of as the
difference. Then this can be written as an expectation value of a relative entropy. 19
S − −S(x) x∼Q G ∗ (x)
=
d
J ∗ (d|x) log Q J ∗ (d|x) x∼Q G ∗ (x)
Transform this into an integration
− Q(d)
(6.104)
log Q(d)
=
d
dx Q G ∗ (x)Q J ∗ (d|x) log Q J ∗ (d|x) −
dx Q G ∗ (x)Q J ∗ (d|x) log Q(d)
=
dx Q G ∗ (x)
d
Q J ∗ (d|x) log
Q J ∗ (d|x)
Q(d)
=
dx Q G ∗ (x)D KL
Q J ∗ (d|x)
Q(d)
=
D KL
Q J ∗ (d|x)
Q(d)
x∼Q G ∗ (x)
.
(6.109)
19 Yet another transform
(6.109) =
dx
d
Q J ∗ G ∗ (x, d) log
Q J ∗ G ∗ (x, d)
Q G ∗ (x)Q(d)
, Q J ∗ G ∗ (x, d) = Q J ∗ (d|x)Q G ∗ (x)
(6.108)
provides a quantity called mutual information between the generated image and the classification
labels. Reference [93] used this to hack IS, that is, generate images that have an unusually large
value of the IS (although images that make little sense to the human eyes). As seen from this
example, it is not always true that the higher the IS, the better.
