116
6 Unsupervised Deep Learning
For the same reason, we find
0 =
1 P (x)>Q G ∗ (x) dx,
(6.73)
and now, almost everywhere, P (x) ≤ Q G ∗ (x). So after all, almost everywhere
P (x) = Q G ∗ (x) .
(6.74)
This is what we wanted to show.
6.3.2 Wasserstein GAN
Another interesting extension [82] of the GAN is possible by considering unsupervised learning in the context of optimal transport [83, 84]. Optimal transport
treats minimization of transport energy, so the model can be interpreted physically.
Regarding the probability distribution Q as a pile of sand, and the probability
distribution P as a hole of the same volume dug in the ground, then the optimal
transport energy is defined as the minimum energy consumed for transporting sand
from the pile to the hole to make a flat surface. One of the typical optimal transport
energies is what is called the Wasserstein distance. The definition is
D W (P , Q) = min
π∈ (6.77)
U(π),
(6.75)
U(π) = =E(x, y) (x,y)∼π(x,y) .
(6.76)
Here E(x, y) represents the transportation cost energy between data points x, y.
Typically, we take the distance between x, y for it. π(x, y) is a joint probability
distribution of how much to transport y → x, which is assumed to satisfy
dx π(x, y) = Q(y),
dy π(x, y) = P (x) .
(6.77)
U(π) is the expectation value of the transportation cost energy using this probability
distribution π, and in the analogy to thermodynamics we can think of it as “internal
energy.” If the transportation cost E(x, y) satisfies the property of distance:
0 ≤ D W (P , Q),
0 = D W (P , Q) ⇔ ∀x, P (x) = Q(x),
(6.78)
it has properties as a substitute for the relative entropy.
Précédent

- 124/211

Suivant