192
F. Nielsen and K. Sun
(KLD also called relative entropy) between parametric models yields practical learning machines. However, the KLD or in general the Csiszár’s f -divergences between
statistical mixtures [44] do not admit closed-form formula, and needs in practice to
be approximated by costly Monte Carlo stochastic integration. To tackle this computational tractability problem, two research directions have been considered in the
literature: ➀ propose some new distances between mixtures that yield closed-form
formula [36, 38] (e.g., the Cauchy-Schwarz divergence, the Jensen quadratic Rényi
divergence, the statistical Minkowski distances). ➁ lower and/or upper bound the
f -divergences between mixtures [15, 44]. However, this direction is tricky when
considering bounded divergences like the Total Variation (TV) distance or the Jensen–
Shannon (JS) divergence that are upper bounded by 1 and log 2, respectively, or when
considering high-dimensional mixtures.
When dealing with probability densities, two main classes of statistical distances
have been widely studied in the literature: ➀ The invariant f -divergences of Information Geometry ([2]; IG) characterized as the class of separable distances which are
information monotone (i.e., satisfies the partition inequality [63]), and ➁ The Optimal
Transport (OT)/Wasserstein/EMD distance [34, 56] which can be computationally
accelerated using entropy regularization [8, 17] (i.e., the Sinkhorn divergence).
In general, computing closed-form formula for the OT between parametric distributions is difficult except in 1D [51]. A closed-form formula is known for elliptical
distributions [13] for the 2-Wasserstein metric (including the multivariate Gaussian
distributions), and the OT of multivariate continuous distributions can be calculated
from the OT of their copulas [22].
The geometry related to these OT/IG distances are different. For example, consider
univariate location-scale families (or multivariate elliptical distributions): For OT,
the 2-Wasserstein distance between any two members admit the same closed-form
formula [13, 21] (depending only on the mean and variance parameters, and not on
the type of location-scale family). The OT geometry of Gaussian distributions has
positive curvature [20, 60]. For any smooth f -divergence, the information-geometric
manifold has negative curvature ([30]; hyperbolic geometry).
In this chapter, we first generalize the work of [33] that proposed a novel family
of statistical distances between statistical mixtures (that we term MCOTs, standing
for Mixture Component Optimal Transports) by solving linear programs between
mixture component weights where the elementary distance between any two mixture components is prescribed. Then we propose to learn Gaussian mixture models
(GMMs) by simplifying kernel density estimators (KDEs) using our distance.
We describe our main contributions as follows:
• We define the generic Chain Rule Optimal Transport (CROT) distance in Definition 8.1, and prove that the CROT distance is a metric whenever the distance
between conditional distributions is a metric in Property 8.2. The CROT distance
unifies and extends the Wasserstein distances and the MCOT distance [33] between
statistical mixtures.
• We report a novel generic upper bound for statistical distances between marginal
distributions [45] in Sect. 8.3 (Theorem 8.8) whenever the ground distance is jointly
F. Nielsen and K. Sun
(KLD also called relative entropy) between parametric models yields practical learning machines. However, the KLD or in general the Csiszár’s f -divergences between
statistical mixtures [44] do not admit closed-form formula, and needs in practice to
be approximated by costly Monte Carlo stochastic integration. To tackle this computational tractability problem, two research directions have been considered in the
literature: ➀ propose some new distances between mixtures that yield closed-form
formula [36, 38] (e.g., the Cauchy-Schwarz divergence, the Jensen quadratic Rényi
divergence, the statistical Minkowski distances). ➁ lower and/or upper bound the
f -divergences between mixtures [15, 44]. However, this direction is tricky when
considering bounded divergences like the Total Variation (TV) distance or the Jensen–
Shannon (JS) divergence that are upper bounded by 1 and log 2, respectively, or when
considering high-dimensional mixtures.
When dealing with probability densities, two main classes of statistical distances
have been widely studied in the literature: ➀ The invariant f -divergences of Information Geometry ([2]; IG) characterized as the class of separable distances which are
information monotone (i.e., satisfies the partition inequality [63]), and ➁ The Optimal
Transport (OT)/Wasserstein/EMD distance [34, 56] which can be computationally
accelerated using entropy regularization [8, 17] (i.e., the Sinkhorn divergence).
In general, computing closed-form formula for the OT between parametric distributions is difficult except in 1D [51]. A closed-form formula is known for elliptical
distributions [13] for the 2-Wasserstein metric (including the multivariate Gaussian
distributions), and the OT of multivariate continuous distributions can be calculated
from the OT of their copulas [22].
The geometry related to these OT/IG distances are different. For example, consider
univariate location-scale families (or multivariate elliptical distributions): For OT,
the 2-Wasserstein distance between any two members admit the same closed-form
formula [13, 21] (depending only on the mean and variance parameters, and not on
the type of location-scale family). The OT geometry of Gaussian distributions has
positive curvature [20, 60]. For any smooth f -divergence, the information-geometric
manifold has negative curvature ([30]; hyperbolic geometry).
In this chapter, we first generalize the work of [33] that proposed a novel family
of statistical distances between statistical mixtures (that we term MCOTs, standing
for Mixture Component Optimal Transports) by solving linear programs between
mixture component weights where the elementary distance between any two mixture components is prescribed. Then we propose to learn Gaussian mixture models
(GMMs) by simplifying kernel density estimators (KDEs) using our distance.
We describe our main contributions as follows:
• We define the generic Chain Rule Optimal Transport (CROT) distance in Definition 8.1, and prove that the CROT distance is a metric whenever the distance
between conditional distributions is a metric in Property 8.2. The CROT distance
unifies and extends the Wasserstein distances and the MCOT distance [33] between
statistical mixtures.
• We report a novel generic upper bound for statistical distances between marginal
distributions [45] in Sect. 8.3 (Theorem 8.8) whenever the ground distance is jointly
