200
C. Greco and M. Vasile
moments can be computed as in Eqs. (6.48) and (6.49) by weighting each sample
with a measure of the deviation between the original and sampled distribution. The
requirement on the importance density is that its support should be greater or equal
to the one of p(x) [38]. To compute the weight, we can decompose the expectation
formula as [51]:
E{z} = E p {g(x)} =
g(x)p(x)dx
=
g(x)
p(x)
π(x)
π(x)dx = E π
g(x)
p(x)
π(x)
.
(6.50)
Therefore, as the samples x i are generated using π(x), their function evaluation
should be weighted by p(x)/π(x). The approximation of the expected value of z
now becomes:
ˆ
z ≈
1
N
N
i=1
z i =
1
N
N
i=1
p(x i )
π(x i )
g(x i ) =
N
i=1
w i g(x i ) .
(6.51)
Hence, the weight of the ith sample is:
w i =
1
N
p(x i )
π(x i )
.
(6.52)
Intuitively, the weights correct the bias associated to the samples selected from a
nonideal distribution. Clearly, the closer the importance distribution is to the original
one, the smaller the required bias correction is, i.e. w i ≈ 1/N. The same weights
computed for the expected value can be used for the covariance or higher-order
moment approximation. Liu [38] suggests to use the normalised weights:
w
∗
i =
w i
i w i
.
(6.53)
The resulting estimate, although biased, often results in a smaller mean squared
error. The same weight choice is found in Sarkka [51] when deriving the importance
sampling form for conditional probabilities (see Sect. 6.3.5).
Indeed, this framework applies also for the density functions in the sequential
filtering algorithms, where p(x) and π(x) are substituted by conditional distributions. In an attempt to connect the generic notation above to the filtering problem of
interest, the vector x can be decomposed as x = x 0:k = [x 0 , . . . , x k ]. Hence, looking
at Eq. (6.50) with x 0:k and substituting a density conditional on measurements
y 1:k , we obtain the importance sampling approximation to the expectation operator
of the sequential filtering posterior distribution. However, the basic importance
sampling approach is not well-suited for sequential filtering approaches. Indeed,
when computing p(x 0:k |y 1:k ) by the importance distribution π(x 0:k |y 1:k ), it would
C. Greco and M. Vasile
moments can be computed as in Eqs. (6.48) and (6.49) by weighting each sample
with a measure of the deviation between the original and sampled distribution. The
requirement on the importance density is that its support should be greater or equal
to the one of p(x) [38]. To compute the weight, we can decompose the expectation
formula as [51]:
E{z} = E p {g(x)} =
g(x)p(x)dx
=
g(x)
p(x)
π(x)
π(x)dx = E π
g(x)
p(x)
π(x)
.
(6.50)
Therefore, as the samples x i are generated using π(x), their function evaluation
should be weighted by p(x)/π(x). The approximation of the expected value of z
now becomes:
ˆ
z ≈
1
N
N
i=1
z i =
1
N
N
i=1
p(x i )
π(x i )
g(x i ) =
N
i=1
w i g(x i ) .
(6.51)
Hence, the weight of the ith sample is:
w i =
1
N
p(x i )
π(x i )
.
(6.52)
Intuitively, the weights correct the bias associated to the samples selected from a
nonideal distribution. Clearly, the closer the importance distribution is to the original
one, the smaller the required bias correction is, i.e. w i ≈ 1/N. The same weights
computed for the expected value can be used for the covariance or higher-order
moment approximation. Liu [38] suggests to use the normalised weights:
w
∗
i =
w i
i w i
.
(6.53)
The resulting estimate, although biased, often results in a smaller mean squared
error. The same weight choice is found in Sarkka [51] when deriving the importance
sampling form for conditional probabilities (see Sect. 6.3.5).
Indeed, this framework applies also for the density functions in the sequential
filtering algorithms, where p(x) and π(x) are substituted by conditional distributions. In an attempt to connect the generic notation above to the filtering problem of
interest, the vector x can be decomposed as x = x 0:k = [x 0 , . . . , x k ]. Hence, looking
at Eq. (6.50) with x 0:k and substituting a density conditional on measurements
y 1:k , we obtain the importance sampling approximation to the expectation operator
of the sequential filtering posterior distribution. However, the basic importance
sampling approach is not well-suited for sequential filtering approaches. Indeed,
when computing p(x 0:k |y 1:k ) by the importance distribution π(x 0:k |y 1:k ), it would
