8 Chain Rule Optimal Transport
203
The OT distance W 2 between Gaussian measures [13, 60] is available in closedform:
W 2 (N (μ 1 , , 1 ), N (μ 1 , , 1 )) =
μ 1 − μ 2 2 + tr(( 1 + 2 − 2((
1
2
1 2
1
2
1 )
1
2 ).
This H
1
p
W
p
p
CROT distance generalizes [7] who considered the W 2 distance
between GMMs using discrete OT. They proved that H W 2 (m 1 , m 2 ) is a metric, and
W 2 (m 1 , m 2 ) ≤
H W
2
2
(m 1 , m 2 ). These results generalize to mixture of elliptical distributions [13]. However, we do not know a closed-form formula for W p between
Gaussian measures when p = 2.
Given two high-dimensional mixture models m 1 and m 2 , we draw respectively
n i.i.d. samples from m 1 and m 2 , so that m 1 (x) ≈
1
n
n
i=1 D(x i ) and m 2 (x) ≈
1
n
n
j=1 D(y j ). Then, we have
W p (m 1 , m 2 ) ≈ W p
⎛
⎝ 1
n
n
i=1
D(x i ),
1
n
n
j=1
D(y j )
⎞
⎠
≤ H
1/ p
W
p
p
⎛
⎝ 1
n
n
i=1
D(x i ),
1
n
n
j=1
D(y j )
⎞
⎠ .
(8.12)
Note that W p
D(x i ), D(x j )
= =x i − x j 2 and therefore the RHS of (8.12) can be
evaluated. We use UB(W 2 ) to denote this empirical upper bound that will hold if
n → ∞. In our experiments n = 10
3 .
See Table 8.2 for the W 2 distances evaluated on the two investigated data sets.
The column LB(W 2 ) is a lower bound based on the first and second moments of the
mixture models [21]. We can clearly see that
H W
2
2
provides a tighter upper bound
than UB(W 2 ). To compute UB(W 2 ) one need to draw a potentially large number of
random samples to make the approximation in (8.12), and the computation of the
EMD is costly. Therefore one should use
H W
2
2
for its better and more efficient
approximation.
8.4.3 Rényi CROT Between GMMs
We investigate Rényi α-divergence [40, 41] defined by
R α ( p : q) =
1
1 − α
log
p(x)
α q(x)
1−α dx,
Précédent

- 212/282

Suivant