9.5 Iteratively Reweighted Least-Square Method
245
9.5 Iteratively Reweighted Least-Square Method
To easily determine correct weights, it is tempting to use the iteratively reweighted
least-squares method (Hamilton 1992; Watson 2003). When the number of data is
large, this method gives good results (Rudolph et al. 2013). However, when the
number of data is small, (which is a frequent case) and is not always trustworthy.
Furthermore, it sometimes unbalances the fit, some data having very high weights
and others very low weights. The formula proposed by Watson (2003) to calculate
the weights often gives satisfactory results,
W i =
1
σ
2
i + r
2
i /3
(9.34)
where r i is the residual of the previous iteration and σ i the estimated uncertainty of
the ith data. If possible, the weight finding should be applied separately to each of
the component sets g = a, b, c of the inertial moments I g as well as to the other sets
of data.
When the number of data is large enough and when the uncertainties of each set
of data are similar, a more sophisticated treatment may be used (Hamilton 1992).
The first step is performed with the standard least-squares method (with no
weighting or with an approximate weighting), then, the weight matrix is updated
to some non-negative function g(r) of the residuals r. With these new weights, the
weighted least-squares equation is solved. The process is iterated in the following
way:
Step 0. Initial residuals r i and leverage h ii are calculated using the ordinary leastsquares method.
Step 1. The median
1 of absolute deviations (MAD) of the residuals is calculated
MAD = median|r i −median(r i )|
(9.35)
and the standard deviation is estimated from MAD as
s = 1.4826
1 +
5
n − p
MAD
(9.36)
(this estimate of the standard deviation is more robust than the one obtained with
(9.6)).
The residuals are scaled by dividing them by s from (9.36)
1 If the n data is listed in order of increasing magnitude, the median is the value at position (n +
1)/2. If n is odd, the median is the middle value, and if n is even, it is the mean of the two middle
values.
Précédent

- 260/291

Suivant