6.4 Generalization in Generative Models
119
Then, adopting f = D, we find that this is a minimax problem (6.41) in the zerosum GAN! Namely, for
V D (G, D) = =D(x) x∼P (x) − −D(y) y∼Q(y) ,
(6.90)
V G (G, D) = −V D (G, D) ,
(6.91)
the equilibrium point of the GAN is also expected from the property of D W (P , Q G )
to satisfy
P (x) = Q G (x) .
(6.92)
One should note that here D must be Lipschitz-continuous. When making D
with a neural network, normally this condition is not satisfied. The original paper
implements an approximate Lipschitz continuity by restricting the range of the
values of the weights. And later years there appeared some ideas such as an
approximate implementation with a gradient penalty (||∇ x D(x)|| 2 − 1) 2 for D as a
regularization term [86], or a spectral normalization in which the network weights
are normalized by their maximum singular value [87]. 17 The gradient penalty and
the spectral normalization are also known to improve the performance of other
GANs which are not WGAN.
6.4 Generalization in Generative Models
So far, we have introduced various models and algorithms of unsupervised machine
learning and deep learning. 18 All of those are aimed to make
Q J ∗ (x) = P (x) .
(6.93)
However, as we have emphasized many times, we cannot access the data generation
probability distribution P (x) in machine learning in the first place. So the learning
algorithm has to use the empirical distribution ˆ
P (x) for P (x), and how should we
think about generalization?
17 Actually, spectral normalization was introduced to stabilize the learning process of conventional
GANs rather than to use it for WGANs, and the authors are not aware of successful examples
of using spectral normalization in WGAN implementations. With this normalization, the network
acquires a K-Lipschitz continuity with a certain number K, so there is no reason why it cannot be
used for WGAN.
18 Noteworthy deep generative models that we have not been able to introduce here include
variational auto-encoder (VAE) [88, 89] and nonlinear independent component estimation (NICE)
[90]. A brief review is provided by a physicist, L. Wang [91].
Précédent

- 127/211

Suivant