194
Chapter 10: The Evaluation of Forecasts
A warning is in order heJ;e: There are at least two other simple ways to inflate ß that usually do not imply any meaningful gain in forecast information
content. Inspection of (10.9) and (10.10), for instance, reveals that general
smoothing of less skillful forecasts, Le., ones with small anomaly correlations
relative to the ratio of forecast to observation map standard deviations or
average anomalies that are greater in magnitude or opposite in sign to the
average observational anomaly, leads to reductions in B2 and C 2 respectively.
Likewise, it is dear from (10.6), the definition of the Brier-based score, that
the use of some inappropriate dimatologies can lead to an increase in M S Ee:c
and thereby an increase of ß. This can occur if estimates of the dimatology
are based on a sampie that is either too small in size (generally it should be
around 30 years) or substantially out of date. The problem is exacerbated
when interdecadal variability is present in the data, as is often the case in
dimate studies.
With these cautions in mind, it will turn out in many applications that the
separation of conditional and unconditional forecast biases from the skill in
phase prediction in (10.10) will provide useful diagnostic information about
the fore cast system which may lead to its improvement. In any event, it is
dear that a multi-faceted perspective like the Murphy-Epstein decomposition
is necessary to properly assess pattern prediction skill.
10.5 Hindcasts and Cross-Validation
Often the skill of a statistical climate prediction model is estimated on the
sampie used to develop the model. This is necessitated by a sampie size that
is too small to support both approximation ofmodel parameters and accurate
estimation of performance on data independent of the subsampie on which
the parameter estimation was based. Such a dependent-data based estimate
is referred to as hindcast skill and is an overestimate of skill that is expected
for future forecasts. The difference between hindcast and true forecast skill
is called artificial skilI.
There are two notable sources of artificial skill. The first is non-stationarity
of the dimate over the time scale spanning the developmental sampie period
and the expected period of application of the statistical forecast model. If
it cannot be reasonably assumed that this non-stationarity is unimportant,
then a source of error is present for independent forecasts that is not present
du ring the developmental period. The assumption of stationarity is usually
reasonable for prediction of interannual variability.
Artificial skill also arises in the statistical modeling process because fitting
of some noise in the dependent sam pIe is unavoidable. The ideal situation
of course is to have a large enough sampIe that this overfitting is minimized,
but this is usually not the case for the same reasons that reservation of an
independent testing sampIe is often not feasible.
There are two approaches available to deal with small sampie estimation
Chapter 10: The Evaluation of Forecasts
A warning is in order heJ;e: There are at least two other simple ways to inflate ß that usually do not imply any meaningful gain in forecast information
content. Inspection of (10.9) and (10.10), for instance, reveals that general
smoothing of less skillful forecasts, Le., ones with small anomaly correlations
relative to the ratio of forecast to observation map standard deviations or
average anomalies that are greater in magnitude or opposite in sign to the
average observational anomaly, leads to reductions in B2 and C 2 respectively.
Likewise, it is dear from (10.6), the definition of the Brier-based score, that
the use of some inappropriate dimatologies can lead to an increase in M S Ee:c
and thereby an increase of ß. This can occur if estimates of the dimatology
are based on a sampie that is either too small in size (generally it should be
around 30 years) or substantially out of date. The problem is exacerbated
when interdecadal variability is present in the data, as is often the case in
dimate studies.
With these cautions in mind, it will turn out in many applications that the
separation of conditional and unconditional forecast biases from the skill in
phase prediction in (10.10) will provide useful diagnostic information about
the fore cast system which may lead to its improvement. In any event, it is
dear that a multi-faceted perspective like the Murphy-Epstein decomposition
is necessary to properly assess pattern prediction skill.
10.5 Hindcasts and Cross-Validation
Often the skill of a statistical climate prediction model is estimated on the
sampie used to develop the model. This is necessitated by a sampie size that
is too small to support both approximation ofmodel parameters and accurate
estimation of performance on data independent of the subsampie on which
the parameter estimation was based. Such a dependent-data based estimate
is referred to as hindcast skill and is an overestimate of skill that is expected
for future forecasts. The difference between hindcast and true forecast skill
is called artificial skilI.
There are two notable sources of artificial skill. The first is non-stationarity
of the dimate over the time scale spanning the developmental sampie period
and the expected period of application of the statistical forecast model. If
it cannot be reasonably assumed that this non-stationarity is unimportant,
then a source of error is present for independent forecasts that is not present
du ring the developmental period. The assumption of stationarity is usually
reasonable for prediction of interannual variability.
Artificial skill also arises in the statistical modeling process because fitting
of some noise in the dependent sam pIe is unavoidable. The ideal situation
of course is to have a large enough sampIe that this overfitting is minimized,
but this is usually not the case for the same reasons that reservation of an
independent testing sampIe is often not feasible.
There are two approaches available to deal with small sampie estimation
