6 Fundamentals of Filtering
193
The optimal estimate of this estimator with W being the identity matrix is the
conditional expectation ˆ
x ∗
k = E{x k |y 1:k } of the posterior distribution. In linear
filtering, under Gaussian assumptions, this estimate coincides with the conditional
mode and median. This state estimate is also called minimum variance estimate,
because it minimises the variance for any conditional probability, or minimum mean
squared error, as it can be derived by minimising the least square errors between
computed and received observations [58].
Still Cox introduced a loss function that weights deviations larger than a set
threshold equally in order to avoid few very dispersed samples to spoil the estimate,
which in this framework tends to the mode of the conditional distribution.
L( x k ) =
|| x k || 2 if || x k || 2 ≤ c
a
if || x k || 2 > c .
(6.30)
The mode of the conditional distribution is exactly the state best estimate when
the Dirac’s delta δ(·) is used in the loss function [37]:
L( x k ) = 1 − δ( x k ) =
0 if x k = ˆ
x k
1 if x k = ˆ
x k .
(6.31)
This loss function choice yields the maximum a posteriori (MAP) estimator as the
minimising point is the peak of the posterior distribution. This can be seen as a
particular case of the function in Eq. (6.30), when no loss is associated to the correct
point and equal loss to any deviation from it.
For the scalar case, Kalman [36] introduced a quartic loss function L( x k ) = a x 4
k
and an exponential one L( x k ) = a[1 − exp
− x 2
k
].
If p(x k |y 1:k ) is symmetric about its conditional expectation and unimodal, the
optimal state estimation is the conditional expectation ˆ
x ∗
k = E{x k |y 1:k } for every
loss function L satisfying the properties in Eq. (6.26) [29]. Therefore the conditional
expectation is the chosen estimate for a large variety of filtering problems.
During this theoretical development, the process of deriving an optimal estimate
relied on the assumption that a full posterior distribution p(x k |y 1:k ) would be
available. However, in the general nonlinear filtering case, it is often impracticable
to derive the full posterior distribution. To deal with this issue, practical methods
to efficiently compute the first moments of p(x k |y 1:k ) have been developed, and
they will be presented in later sections. Hence, if the state estimate is chosen as
the conditional expectation, no further calculation is needed as ˆ
x ∗
k coincides with
the first moment of the posterior distribution. On the other hand, we cannot just
look for the conditional mean. First, it often depends on higher-order moments.
Second, higher-order moments are an indication of how the probability is dispersed
around the mean value, therefore providing a measure of how accurate the estimate
represents the distribution. In the words of Jazwinski: ‘It can be argued that
knowledge of the second-order moment is just as important as knowing the estimate
itself. An estimate is meaningless unless one knows how good it is.’
193
The optimal estimate of this estimator with W being the identity matrix is the
conditional expectation ˆ
x ∗
k = E{x k |y 1:k } of the posterior distribution. In linear
filtering, under Gaussian assumptions, this estimate coincides with the conditional
mode and median. This state estimate is also called minimum variance estimate,
because it minimises the variance for any conditional probability, or minimum mean
squared error, as it can be derived by minimising the least square errors between
computed and received observations [58].
Still Cox introduced a loss function that weights deviations larger than a set
threshold equally in order to avoid few very dispersed samples to spoil the estimate,
which in this framework tends to the mode of the conditional distribution.
L( x k ) =
|| x k || 2 if || x k || 2 ≤ c
a
if || x k || 2 > c .
(6.30)
The mode of the conditional distribution is exactly the state best estimate when
the Dirac’s delta δ(·) is used in the loss function [37]:
L( x k ) = 1 − δ( x k ) =
0 if x k = ˆ
x k
1 if x k = ˆ
x k .
(6.31)
This loss function choice yields the maximum a posteriori (MAP) estimator as the
minimising point is the peak of the posterior distribution. This can be seen as a
particular case of the function in Eq. (6.30), when no loss is associated to the correct
point and equal loss to any deviation from it.
For the scalar case, Kalman [36] introduced a quartic loss function L( x k ) = a x 4
k
and an exponential one L( x k ) = a[1 − exp
− x 2
k
].
If p(x k |y 1:k ) is symmetric about its conditional expectation and unimodal, the
optimal state estimation is the conditional expectation ˆ
x ∗
k = E{x k |y 1:k } for every
loss function L satisfying the properties in Eq. (6.26) [29]. Therefore the conditional
expectation is the chosen estimate for a large variety of filtering problems.
During this theoretical development, the process of deriving an optimal estimate
relied on the assumption that a full posterior distribution p(x k |y 1:k ) would be
available. However, in the general nonlinear filtering case, it is often impracticable
to derive the full posterior distribution. To deal with this issue, practical methods
to efficiently compute the first moments of p(x k |y 1:k ) have been developed, and
they will be presented in later sections. Hence, if the state estimate is chosen as
the conditional expectation, no further calculation is needed as ˆ
x ∗
k coincides with
the first moment of the posterior distribution. On the other hand, we cannot just
look for the conditional mean. First, it often depends on higher-order moments.
Second, higher-order moments are an indication of how the probability is dispersed
around the mean value, therefore providing a measure of how accurate the estimate
represents the distribution. In the words of Jazwinski: ‘It can be argued that
knowledge of the second-order moment is just as important as knowing the estimate
itself. An estimate is meaningless unless one knows how good it is.’
