120
6 Unsupervised Deep Learning
Case of rote memorization
First, if we simply replace (6.93) with the empirical distribution,
Q J ∗ (x) = ˆ
P (x) ,
(6.94)
it simply appears that “the machine remembers the data that came out.” Actually,
we want the machine to create new images that do not exist in the data, so this
is not sufficient. However, there are things we can learn from this setup. We
have already derived an inequality of generalization performance in unsupervised
machine learning of the rote memorization: it is (5.21) which was explained at
the end of the section for the central limit theorem. Here we repeat explaining
the problem again. An event with an occurrence probability of p 1 , p 2 , . . . , p W
is defined as A 1 , A 2 , . . . , A W , and the specific value of the probability p i is not
known. Instead, we have only the data
•
The number of event A 1 is # 1 , the number of A 2 is # 2 , . . . , the number of A W is # W ,
In total, the number of events is # =
W
i=1 # i .
(6.95)
and the problem is to estimate the value of p i . This problem can be thought of as
an unsupervised learning. Obviously, the appropriate estimated probability is the
“empirical distribution” q i =
# i
# , and if this probability is close to p i , it can be said
that the model is generalized. The central limit theorem says
With approximately 70% probability,
one finds p i −
p i (1 − p i )
#
< q i < p i +
p i (1 − p i )
#
.
(6.96)
Therefore, even in the case of the unsupervised learning, the generalization performance is expected to be of the order of 1/
√
#.
Actual cases
Do actual deep generative models memorize everything by rote? In the case of
using the actual empirical distribution ˆ
P (x), the proof for (6.93) of GANs can be
applied exactly as it is, and it appears that (6.94) is proven. However, there are other
differences between the proof and the actual cases:
• The convergence proof of GANs assumes that G and D have infinite expressive
power, while in the actual training, G and D are deep neural networks with a
fixed structure, and their expression is limited.
• The Nash equilibrium condition (6.39), (6.40) which is assumed in the proof
is actually solved by the gradient method (6.37), (6.38). However, the gradient
method is not exact, and therefore, the obtained trained model is not an exact
equilibrium point.
Précédent

- 128/211

Suivant