6.4 Generalization in Generative Models
121
Fig. 6.1 Left: Images created by G. Right: Image created by G (upper left red frame) and 9 images
closest to it in the training data
Therefore, it is too early to say that the proof of (6.93) under ideal assumptions is
actually “only” (6.94), which is a replacement with an empirical distribution. Then,
what is the conclusion, in the end? A picture is worth a thousand words. Here, we
show the image generated by a GAN which was actually trained with CIFAR-10, in
Fig. 6.1 (left). With this alone, we cannot tell more than “something that looks like
a natural image is generated,” so in Fig. 6.1 (right) we show the image generated by
the GAN and the training data in the neighborhood of it found in the image space.
A glance finds no image which looks exactly the same as the image generated by G,
and furthermore, even in the closest image the detailed structure is different. This is
supporting evidence that the GAN, a kind of deep generative model, avoids a rote
memorization (6.94) and generalizes.
Inception score
Let us introduce an attempt to measure a “generalization performance” in
image generation from a completely different perspective. In the first place, we
regard (6.93) as a state which generalizes, because the sampling from this generator
gives us a “realistic” image x. Hence the generalization performance is whether
images that “look real to humans” can be generated or not. On the other hand, many
“image classification networks” led by ResNet [24] trained using the ImageNet
dataset [26] are now known to have classification accuracy superior to that of
humans [92]. So the idea is to have the image classification network decide whether
the generated image is “real-looking to humans” or not. Let us prepare
Probability distribution given by generator: Q G ∗ (x),
(6.97)
Trained image classification network: Q J ∗ (d|x).
(6.98)
121
Fig. 6.1 Left: Images created by G. Right: Image created by G (upper left red frame) and 9 images
closest to it in the training data
Therefore, it is too early to say that the proof of (6.93) under ideal assumptions is
actually “only” (6.94), which is a replacement with an empirical distribution. Then,
what is the conclusion, in the end? A picture is worth a thousand words. Here, we
show the image generated by a GAN which was actually trained with CIFAR-10, in
Fig. 6.1 (left). With this alone, we cannot tell more than “something that looks like
a natural image is generated,” so in Fig. 6.1 (right) we show the image generated by
the GAN and the training data in the neighborhood of it found in the image space.
A glance finds no image which looks exactly the same as the image generated by G,
and furthermore, even in the closest image the detailed structure is different. This is
supporting evidence that the GAN, a kind of deep generative model, avoids a rote
memorization (6.94) and generalizes.
Inception score
Let us introduce an attempt to measure a “generalization performance” in
image generation from a completely different perspective. In the first place, we
regard (6.93) as a state which generalizes, because the sampling from this generator
gives us a “realistic” image x. Hence the generalization performance is whether
images that “look real to humans” can be generated or not. On the other hand, many
“image classification networks” led by ResNet [24] trained using the ImageNet
dataset [26] are now known to have classification accuracy superior to that of
humans [92]. So the idea is to have the image classification network decide whether
the generated image is “real-looking to humans” or not. Let us prepare
Probability distribution given by generator: Q G ∗ (x),
(6.97)
Trained image classification network: Q J ∗ (d|x).
(6.98)
