110
6 Unsupervised Deep Learning
Objective function and learning process
The interesting point about GANs is that it does not use the relative entropy (6.3)
directly to optimize (6.32). Instead, one considers the following function: 6
V D (G, D) = =log(1 + e
−D(x) ) x∼P (x) + +log(1 + e
+D(x) ) x∼Q G (x) ,
(6.35)
V G (G, D) = −V D (G, D).
(6.36)
This just means: (the value of V D is small) ⇔ (D(x real ) =large, D(x fake ) = small).
If we train D to reduce V D , then we can put a role as a “police officer” to D. On
the other hand, according to (6.36), decreasing V G = increasing V D , so if this value
of V G is reduced, G can trick D. Namely, the learning process of GANs is to repeat
the following updates:
G ← G − G V G (G, D),
(6.37)
D ← D − G V D (G, D),
(6.38)
This optimization is a search for a Nash equilibrium,
G
∗ s.t.∀G, V G (G
∗ , D
∗ ) ≤ V G (G, D
∗ ),
(6.39)
D
∗ s.t.∀D, V D (G
∗ , D
∗ ) ≤ V D (G
∗ , D).
(6.40)
In fact, with the condition of (6.36), that is, the sum of the objective function of G
and the objective function of D becomes zero, the equilibrium point is shown to
satisfy
V D (G
∗ , D
∗ ) = min
G
max
D
V D (G, D).
(6.41)
This is a consequence of the minimax theorem [80] by von Neumann. 7 From this
condition, it is relatively easy to prove why Q G reproduces P . First, we look at
max D of (6.41). D is a function on X and the functional V D is differentiable; this
6 This is written with the sigmoid function σ (u) = (1 + e −u ) −1 as
−−log σ (D(x)) x∼P (x) − −log(1 − σ (D(x))) x∼Q G (x)
(6.34)
and it is similar to the cross-entropy error (3.16), which is the error function derived for binary
classification in Chap. 3. In fact, this is identical to a binary classification problem of discriminating
whether an item of data is real (x ∼ P (x)) or not (x ∼ Q G (x)).
7 The minimax theorem is that if f (x, y) is a concave (convex) function with respect to the 1st
(2nd) variable x (y),
min
x
max
y
f (x, y) = max
y
min
x
f (x, y) .
(6.42)
Taking f = V D and setting the solution of this minimax problem to G, D, one can prove that
these satisfy the Nash equilibrium condition (6.39) and (6.40). In the proof, one uses the condition
V G + V D = 0 (the zero-sum condition).
Précédent

- 118/211

Suivant