276
Chapter 14: Patterns in Time: SSA and MSSA
3. Specification: The first 15 ST-PCs are used to estimate conditional probabilities of occurrence within each tercile. This is done by seeking in the
past for the 200 best analog ST-PCs ofthe forecast ST-PCs and counting
the number of analogs associated with each tercile. For the validation
of the experiment, we choose as a deterministic forecast the tercile that
has the highest conditional prob ability.
This procedure is a poor-man's version of the long-range empirical forecast
scheme used at the U.K. Meteorological office (Folland et al., 1986b), where
first a weather type is forecast (among 6 clusters), and then the specification
procedure in order to determine categories is performed using discriminant
analysis and other more subtle techniques.
The forecasting scheme presented here is validated using a cross-validation
technique described in Chapter 10. Verification data are gathered in 14
groups of 3 years, each being removed from the total dataset in order to
form a learning period from which data can be used to construct the models.
Given a 3-year verification period, the coefficients of the forecasting model
are calculated from data in the associated learning period, and the fore cast
first 15 ST-PCs are compared with all the observed first 15 ST-PCs occurring at each pentad of the learning set, using the eosine of the angle between
the two (pattern correlation) as a similarity measure. The 200 best learning
analogs are extracted. At each Atlantic grid point, their associated monthly
mean tercile is calculated. Thus, the conditional probability of occurrence
of a tercile is the percentage of these 200 analogs falling within the tercile.
Next, for validation purposes, a deterministic forecast is buHt by choosing
the tercile that has the highest conditional prob ability. Verification scores
are stored in a 3 x 3 contingency table, where the diagonal counts the number of successful forecasts, at each grid point. This table is used to calculate
the LEPS categorical score (see Chapter 10) of the hindcast experiment.
Figure 14.5 shows the LEPS score counted over all Atlantic grid points for
the 14 periods of the cross-validation experiment. The scores of "persistence"
forecasts are also represented. Persistence forecasts consist in forecasting, for
the forthcoming month, the same tercile as the one of the past month. The
MSSA scores are generally better than persistence scores which are themselves better than random forecast scores, except in the earliest periods. The
scores lie, in general, around 0.2, except for the latest periods, which seem
to have exceptionally good scores, possibly due to the improvements in the
data themselves, leading to better initial conditions of the forecast.
Figure 14.6 shows the spatial distribution of the average score over the
whole cross-validation period. It is noteworthy that this distribution is not
homogeneous at an, with two areas of good scores near Greenland and off
the European western coasts. This reflects the fact that one pattern is particularly wen forecast: The NAO pattern. Half-way between the two centers
of action of this pattern, there is a band of relatively poor forecasts. NAO is
the most persistent feature of the Atlantic area, and is also strongly involved
in the 70-day oscillation. Yet, notice that the score is everywhere positive.
Chapter 14: Patterns in Time: SSA and MSSA
3. Specification: The first 15 ST-PCs are used to estimate conditional probabilities of occurrence within each tercile. This is done by seeking in the
past for the 200 best analog ST-PCs ofthe forecast ST-PCs and counting
the number of analogs associated with each tercile. For the validation
of the experiment, we choose as a deterministic forecast the tercile that
has the highest conditional prob ability.
This procedure is a poor-man's version of the long-range empirical forecast
scheme used at the U.K. Meteorological office (Folland et al., 1986b), where
first a weather type is forecast (among 6 clusters), and then the specification
procedure in order to determine categories is performed using discriminant
analysis and other more subtle techniques.
The forecasting scheme presented here is validated using a cross-validation
technique described in Chapter 10. Verification data are gathered in 14
groups of 3 years, each being removed from the total dataset in order to
form a learning period from which data can be used to construct the models.
Given a 3-year verification period, the coefficients of the forecasting model
are calculated from data in the associated learning period, and the fore cast
first 15 ST-PCs are compared with all the observed first 15 ST-PCs occurring at each pentad of the learning set, using the eosine of the angle between
the two (pattern correlation) as a similarity measure. The 200 best learning
analogs are extracted. At each Atlantic grid point, their associated monthly
mean tercile is calculated. Thus, the conditional probability of occurrence
of a tercile is the percentage of these 200 analogs falling within the tercile.
Next, for validation purposes, a deterministic forecast is buHt by choosing
the tercile that has the highest conditional prob ability. Verification scores
are stored in a 3 x 3 contingency table, where the diagonal counts the number of successful forecasts, at each grid point. This table is used to calculate
the LEPS categorical score (see Chapter 10) of the hindcast experiment.
Figure 14.5 shows the LEPS score counted over all Atlantic grid points for
the 14 periods of the cross-validation experiment. The scores of "persistence"
forecasts are also represented. Persistence forecasts consist in forecasting, for
the forthcoming month, the same tercile as the one of the past month. The
MSSA scores are generally better than persistence scores which are themselves better than random forecast scores, except in the earliest periods. The
scores lie, in general, around 0.2, except for the latest periods, which seem
to have exceptionally good scores, possibly due to the improvements in the
data themselves, leading to better initial conditions of the forecast.
Figure 14.6 shows the spatial distribution of the average score over the
whole cross-validation period. It is noteworthy that this distribution is not
homogeneous at an, with two areas of good scores near Greenland and off
the European western coasts. This reflects the fact that one pattern is particularly wen forecast: The NAO pattern. Half-way between the two centers
of action of this pattern, there is a band of relatively poor forecasts. NAO is
the most persistent feature of the Atlantic area, and is also strongly involved
in the 70-day oscillation. Yet, notice that the score is everywhere positive.
