Section 10.5: Hindcasts and Cross-Validation
195
of true skill. Both require that the climate is negligibly non-stationary on
the time scales of the particular forecast problem. The first is to attempt
to estimate artificial skill simultaneously with hindcast skill, either through
analytical or Monte Carlo means. Techniques for some problems have been
presented in the literature. In particular, a thorough discussion of the considerations involved in artificial skill estimation for linear regression modeling (including screening regression) can be found in the series of papers by
Lanzante (1984), Shapiro (1984), Shapiro and Chelton (1986), and Lanzante
(1986). Much of their thrust is summarized by Michaelson (1987), who simultaneously examines the powerful alternative of cross-validation, the subject
of the remainder of this section.
Cross-validation is arguably the best approach for estimating true skill in
small-sample modelingftesting environments; in most instances it permits
independent verification over the entire available sample with only a small
reduction in developmental sample size. Its key feature is that each fore cast
is based on a different model derived from a subsample of the full data set
that is statistically independent of the forecast. In this way many different
(but highly related) models are used to make independent forecasts. The
next two subsections, which consist of a general description of the procedure
and a discussion of two crucial considerations for application respectively,
should clarify how and why cross-validation works. To augment the material
here, Michaelson (1987) is highly recommended.
10.5.1 Cross-Validation Procedure
For convenience the narratives here and in the next subsection are cast in
terms of a typical interannual prediction problem (for example seasonal prediction) for which predictor and predictand data are matched up by year. The
possibility of other correspondences is understood and easily accommodated.
The cross-validation begins with deletion of one or more years from the
complete sample of available data and construction of a forecast model on
the subsample remaining. Forecasts are then made with this model for one
or more of the deleted years and these forecasts verified.
The deleted years are then returned to the sample and the whole procedure described so far is repeated with deletion of a different group of years.
This new group of deleted years can overlap those previously withheld but
a particular year can only be forecast once during the entire cross-validation
process.
The last step above is repeated until the sample is exhausted and no more
forecasts are possible.
10.5.2 Key Constraints in Cross-Validation
For each iteration of model building and prediction during the course of crossvalidation, every detail of the process must be repeated, including recompu-
195
of true skill. Both require that the climate is negligibly non-stationary on
the time scales of the particular forecast problem. The first is to attempt
to estimate artificial skill simultaneously with hindcast skill, either through
analytical or Monte Carlo means. Techniques for some problems have been
presented in the literature. In particular, a thorough discussion of the considerations involved in artificial skill estimation for linear regression modeling (including screening regression) can be found in the series of papers by
Lanzante (1984), Shapiro (1984), Shapiro and Chelton (1986), and Lanzante
(1986). Much of their thrust is summarized by Michaelson (1987), who simultaneously examines the powerful alternative of cross-validation, the subject
of the remainder of this section.
Cross-validation is arguably the best approach for estimating true skill in
small-sample modelingftesting environments; in most instances it permits
independent verification over the entire available sample with only a small
reduction in developmental sample size. Its key feature is that each fore cast
is based on a different model derived from a subsample of the full data set
that is statistically independent of the forecast. In this way many different
(but highly related) models are used to make independent forecasts. The
next two subsections, which consist of a general description of the procedure
and a discussion of two crucial considerations for application respectively,
should clarify how and why cross-validation works. To augment the material
here, Michaelson (1987) is highly recommended.
10.5.1 Cross-Validation Procedure
For convenience the narratives here and in the next subsection are cast in
terms of a typical interannual prediction problem (for example seasonal prediction) for which predictor and predictand data are matched up by year. The
possibility of other correspondences is understood and easily accommodated.
The cross-validation begins with deletion of one or more years from the
complete sample of available data and construction of a forecast model on
the subsample remaining. Forecasts are then made with this model for one
or more of the deleted years and these forecasts verified.
The deleted years are then returned to the sample and the whole procedure described so far is repeated with deletion of a different group of years.
This new group of deleted years can overlap those previously withheld but
a particular year can only be forecast once during the entire cross-validation
process.
The last step above is repeated until the sample is exhausted and no more
forecasts are possible.
10.5.2 Key Constraints in Cross-Validation
For each iteration of model building and prediction during the course of crossvalidation, every detail of the process must be repeated, including recompu-
