106
F. Barbaresco
5.5 Homogeneous Symplectic Manifold as Co-Adjoint
Orbits and Their Density of Probability: Poincaré Unit
Disk and Its Gaussian Density by Moment Map
of SU(1,1) Lie Group
Classically, to optimize the parameter θ of a probabilistic model, based on a sequence
of observations y t , is an online gradient descent:
θ t ← θ t−1 − η t
∂l t (y t )
T
∂θ
(5.27)
with learning rate η t , and the loss function l t = − log p
y t / ˆ
y t
. This simple gradient
descent has a first drawback of using the same non-adaptive learning rate for all
parameter components, and a second drawback of non invariance with respect to
parameter re-encoding inducing different learning rates. Amari has introduced the
natural gradient to preserve this invariance to be insensitive to the characteristic scale
of each parameter direction. The gradient descent could be corrected by I (θ )
−1 where
I is the Fisher information matrix with respect to parameter θ , given by:
I (θ ) =
g ¯
j
with g i j =
−E y≈ p(y/θ)
∂
2 log p(y/θ )
∂θ i ∂θ j
i j
(5.28)
with natural gradient:
θ t ← θ t−1 − η t I (θ )
−1 ∂l t (y t )
T
∂θ
(5.29)
Amari has proved that the Riemannian metric in an exponential family is the
Fisher information matrix defined by:
g i j = −
∂
2
∂θ i ∂θ j
i j
with = − log
R
e
−−θ,y dy
(5.30)
and the dual potential, the Shannon entropy, is given by the Legendre transform:
S(η) = θ, η − with η i =
∂∂(θ)
∂θ i
and θ t =
∂ S(η)
∂η i
(5.31)
We can observe that = − log
R e
−(θ,y dy = − log ψ(θ) is related to the
classical cumulant generating function.
J. L. Koszul and E. Vinberg have introduced an affinely invariant Hessian metric
on a sharp convex cone through its characteristic function:
F. Barbaresco
5.5 Homogeneous Symplectic Manifold as Co-Adjoint
Orbits and Their Density of Probability: Poincaré Unit
Disk and Its Gaussian Density by Moment Map
of SU(1,1) Lie Group
Classically, to optimize the parameter θ of a probabilistic model, based on a sequence
of observations y t , is an online gradient descent:
θ t ← θ t−1 − η t
∂l t (y t )
T
∂θ
(5.27)
with learning rate η t , and the loss function l t = − log p
y t / ˆ
y t
. This simple gradient
descent has a first drawback of using the same non-adaptive learning rate for all
parameter components, and a second drawback of non invariance with respect to
parameter re-encoding inducing different learning rates. Amari has introduced the
natural gradient to preserve this invariance to be insensitive to the characteristic scale
of each parameter direction. The gradient descent could be corrected by I (θ )
−1 where
I is the Fisher information matrix with respect to parameter θ , given by:
I (θ ) =
g ¯
j
with g i j =
−E y≈ p(y/θ)
∂
2 log p(y/θ )
∂θ i ∂θ j
i j
(5.28)
with natural gradient:
θ t ← θ t−1 − η t I (θ )
−1 ∂l t (y t )
T
∂θ
(5.29)
Amari has proved that the Riemannian metric in an exponential family is the
Fisher information matrix defined by:
g i j = −
∂
2
∂θ i ∂θ j
i j
with = − log
R
e
−−θ,y dy
(5.30)
and the dual potential, the Shannon entropy, is given by the Legendre transform:
S(η) = θ, η − with η i =
∂∂(θ)
∂θ i
and θ t =
∂ S(η)
∂η i
(5.31)
We can observe that = − log
R e
−(θ,y dy = − log ψ(θ) is related to the
classical cumulant generating function.
J. L. Koszul and E. Vinberg have introduced an affinely invariant Hessian metric
on a sharp convex cone through its characteristic function:
