6.3 Generative Adversarial Network
113
As before, one may want to take a variation with respect to D(x), but unfortunately
it does not work because of the condition (6.51) and the discontinuous function
“max.” So, we evaluate what the integrand looks like at each x. Then basically we
just need to consider a one-variable function
f (D) = P D + Q max(0, m − D), 0 < P, Q < 1.
(6.54)
and we find that it is just a matter of division of cases:
min f (D) =
f (D = 0) = mQ (P < Q),
f (D = m) = mP (P ≥ Q).
(6.55)
Then, there is this function for each point x, so we find
V D (G
∗ , D
∗ ) = m
1 P (x) 1 P (x)≥Q G ∗ (x) dx Q G ∗ (x)
.
(6.56)
Here, 1 condition is a step function that takes 1 (0) when the condition is (not) satisfied.
By definition
1 P (x)≥Q G ∗ (x) = 1 − 1 P (x) (6.57)
Substituting this into (6.56), using the fact that Q G ∗ (x) is a probability distribution
and so its integration gives 1, we find
V D (G
∗ , D
∗ ) = m
1 +
1 P (x) P (x) − Q G ∗ (x)
<0
≤ m .
(6.58)
This inequality is the first important conclusion. The point here is that the relationship with the last m is ≤ instead of <. This is the heart of the proof, so let us
elaborate on the explanation. The reason why we do not have < is that there can be
no 10 point x satisfying P (x) − Q G ∗ (x) < 0. In this case, the contribution from the
second term of (6.58) is zero due to the step function 1 P (x) Property obtained from equilibrium point of V G
Next, we write the inequality (6.39) by integration as
∀G,
dx Q G ∗ (x)D
∗ (x) ≤
dx Q G (x)D
∗ (x) .
(6.59)
10 To be precise, it is better to say P (x) − Q G ∗ (x) ≥ 0 almost everywhere. The meaning of this
expression is explained in the next footnote.
Précédent

- 121/211

Suivant