14
G. Pistone
f ∈ L (cosh −1) ( p), that is, f is sub-exponential under the distribution P = p · γ . If
the sequence (X n )
∞
n=1 is independent and with distribution p · γ , then the sequence
of sample means will converge,
lim
n→∞
1
n
n
j=1
f (X j ) =
f (x) p(x) γ (x) dx ,
with an exponential bound on the tail probability. See, for example, [31, Sect. 2.8].
1.4.2 Hyvärinen Divergence
Here we adapt [23] to the Gaussian case. Consider the Hyvärinen divergence of Eq.
(1.1) in the Gaussian case, that is, P = p · γ and Q = q · γ . As a function of q is of
the form
H (q) =
1
2
|∇ log p(x)|
2 p(x) γ (x) dx +
1
2
|∇ log q(x)|
2 p(x) γ (x) dx −
∇ log p(x) · ∇ log q(x) p(x) γ (x) dx ,
where the first term does not depend on q and the second term is an expectation with
respect to p · γ . As ∇ log p = p
−1
∇ p, the third term equals
−
δ · ∇ log q(x) p(x) γ (x) dx ,
which is again a p-expectation. To minimize the Hyvärinen divergence we must
minimize the p-expected value of the local score
S(q, x) =
1
2
|∇ log q(x)|
2
− δ · ∇ log q(x)
If p and q belong to the maximal exponential model of γ , then q = e
u−K (u) with u ∈
L (cosh −1) (γ ) and
u(x) γ (x) dx = 0. The local score becomes
1
2
|∇u|
2
− δ · ∇u.
To compute the p-expected value of the score with an independent sample of p · γ
we have interest to assume that the score is in L (cosh −1) (γ ), because this assumption
implies the good convergence of the empirical means for all p, as it way explained
in the section above.
Assume, for example, ∇u ∈ L
2
(cosh −1) (γ ). This implies directly |∇u|
2
∈
L (cosh −1) (γ ). Moreover, we need to assume that the L (cosh −1) (γ )-norm of δ · ∇u
is finite. Under such assumptions it seems reasonable to hope that the minimization
on a suitable model of the sample expectation of the Hyvärinen score is consistent.
G. Pistone
f ∈ L (cosh −1) ( p), that is, f is sub-exponential under the distribution P = p · γ . If
the sequence (X n )
∞
n=1 is independent and with distribution p · γ , then the sequence
of sample means will converge,
lim
n→∞
1
n
n
j=1
f (X j ) =
f (x) p(x) γ (x) dx ,
with an exponential bound on the tail probability. See, for example, [31, Sect. 2.8].
1.4.2 Hyvärinen Divergence
Here we adapt [23] to the Gaussian case. Consider the Hyvärinen divergence of Eq.
(1.1) in the Gaussian case, that is, P = p · γ and Q = q · γ . As a function of q is of
the form
H (q) =
1
2
|∇ log p(x)|
2 p(x) γ (x) dx +
1
2
|∇ log q(x)|
2 p(x) γ (x) dx −
∇ log p(x) · ∇ log q(x) p(x) γ (x) dx ,
where the first term does not depend on q and the second term is an expectation with
respect to p · γ . As ∇ log p = p
−1
∇ p, the third term equals
−
δ · ∇ log q(x) p(x) γ (x) dx ,
which is again a p-expectation. To minimize the Hyvärinen divergence we must
minimize the p-expected value of the local score
S(q, x) =
1
2
|∇ log q(x)|
2
− δ · ∇ log q(x)
If p and q belong to the maximal exponential model of γ , then q = e
u−K (u) with u ∈
L (cosh −1) (γ ) and
u(x) γ (x) dx = 0. The local score becomes
1
2
|∇u|
2
− δ · ∇u.
To compute the p-expected value of the score with an independent sample of p · γ
we have interest to assume that the score is in L (cosh −1) (γ ), because this assumption
implies the good convergence of the empirical means for all p, as it way explained
in the section above.
Assume, for example, ∇u ∈ L
2
(cosh −1) (γ ). This implies directly |∇u|
2
∈
L (cosh −1) (γ ). Moreover, we need to assume that the L (cosh −1) (γ )-norm of δ · ∇u
is finite. Under such assumptions it seems reasonable to hope that the minimization
on a suitable model of the sample expectation of the Hyvärinen score is consistent.
