184
Chapter 10: The Evaluation of Forecasts
Generally, the control selected should be the one that has the most apparent skill for the measure used in the evaluation. For example, if the measure
used is mean squared error (MSE), random or persistence forecasts should
not be used as no-skill standards because their MSE is quite high compared
to damped persistence or climatology and is easy to improve upon. In fact,
damped persistence is the right choice because it minimizes MSE on the developmental sampie. Properties of commonly-used measures and relationships
between them will be discussed in the next two sections.
When a forecast system is being compared to other competing systems,
apart from the controls a number of additional considerations are important.
First, homogeneo'Us comparisons are preferred, i.e. on the same forecast set,
otherwise the comparisons may be unfair if inhomogeneities are not taken
into account. This is because some situations are easier to predict than
others. Second, none of the forecasts being compared should use data that
are unavailable or arbitrarily withheld from any of the others.
For example, a comparison between a neural net time series model that
looks at antecedent data at four lags and an AR(l) model would be unfair.
The fair comparison would be between the former and higher order autoregressive models that also use up to four lags. Another example is shown in
Figure 10.4 in which monthly forecasters are operationally constrained to produ ce their forecasts several days before the fore cast periods begin. Therefore,
it would be unfair to compare their performance to a persistence forecaster
who utilizes data over entire antecedent months. Instead, "operational" persistence should be used as the control.
Table 10.1: Contingency table for forecasts/observations in three categories.
ForeObserved
cast
B
N
A
B
fBB
fBN
fBA
N
fNB
fNN
fNA
A
fAB
fAN
fAA
Chapter 10: The Evaluation of Forecasts
Generally, the control selected should be the one that has the most apparent skill for the measure used in the evaluation. For example, if the measure
used is mean squared error (MSE), random or persistence forecasts should
not be used as no-skill standards because their MSE is quite high compared
to damped persistence or climatology and is easy to improve upon. In fact,
damped persistence is the right choice because it minimizes MSE on the developmental sampie. Properties of commonly-used measures and relationships
between them will be discussed in the next two sections.
When a forecast system is being compared to other competing systems,
apart from the controls a number of additional considerations are important.
First, homogeneo'Us comparisons are preferred, i.e. on the same forecast set,
otherwise the comparisons may be unfair if inhomogeneities are not taken
into account. This is because some situations are easier to predict than
others. Second, none of the forecasts being compared should use data that
are unavailable or arbitrarily withheld from any of the others.
For example, a comparison between a neural net time series model that
looks at antecedent data at four lags and an AR(l) model would be unfair.
The fair comparison would be between the former and higher order autoregressive models that also use up to four lags. Another example is shown in
Figure 10.4 in which monthly forecasters are operationally constrained to produ ce their forecasts several days before the fore cast periods begin. Therefore,
it would be unfair to compare their performance to a persistence forecaster
who utilizes data over entire antecedent months. Instead, "operational" persistence should be used as the control.
Table 10.1: Contingency table for forecasts/observations in three categories.
ForeObserved
cast
B
N
A
B
fBB
fBN
fBA
N
fNB
fNN
fNA
A
fAB
fAN
fAA
