Chapter 16 . Time-Series Prediction of Marine Zooplankton
325
3
Q)
"0
.a 2
'2
Cl
rtI
E
0
1992
1993
1994
1995
Fig. 16.4. Crossvalidation results for the prediction of Cirripedia nauplius.
Shown are the data and the absolute errors of predictions by 30 Neural Networks.
The structure of the networks was identical, but they were trained by different
data splittings, namely those indicated in Fig. 16.2. The network structure was the
same as in Fig. 16.3.
Nevertheless, a more critical look at the predictions changes this impression. A
first indication that the generalization may not work weIl is the value 0.46 of the
correlation between the data and the prediction error. This value is still significant
so that the predictions cannot be considered to be independent of the data so that
the modelling of the data by the Neural Network is only partially correct. And this
is not the result of an insufficient training: the training was broken off by the early
stopping algorithm at training cycle 364, because the last 100 cycles the validation
error was increasing. The picture gets even worse, if one looks at the fuH
crossvalidation results shown in Fig. 16.4. Obviously there are other splittings of
the training data, for which the error is about 50% of the data maximum, so that
the quality of the predictions depends strongly on the choice of the training data
and the predictions are not reliable. Accordingly, one has to conclude that for this
Neural Network generalization fails.
16.5
Conclusions
In this article we discussed the application of time series prediction by Neural
Networks to environmental data sets. Such data are typically of low quality as
compared to laboratory data for various reasons: boundary conditions for the
phenomenon to be studied are not controHable, taking clean data is much too
expensive or the phenomenon is so complex that the relevant variables that should
be measured are not known. The consequence is: one usually has to work with the
data one can get, and not the data one would like to have. The zooplankton
forecasts considered in the previous section illustrate this situation: the data are
325
3
Q)
"0
.a 2
'2
Cl
rtI
E
0
1992
1993
1994
1995
Fig. 16.4. Crossvalidation results for the prediction of Cirripedia nauplius.
Shown are the data and the absolute errors of predictions by 30 Neural Networks.
The structure of the networks was identical, but they were trained by different
data splittings, namely those indicated in Fig. 16.2. The network structure was the
same as in Fig. 16.3.
Nevertheless, a more critical look at the predictions changes this impression. A
first indication that the generalization may not work weIl is the value 0.46 of the
correlation between the data and the prediction error. This value is still significant
so that the predictions cannot be considered to be independent of the data so that
the modelling of the data by the Neural Network is only partially correct. And this
is not the result of an insufficient training: the training was broken off by the early
stopping algorithm at training cycle 364, because the last 100 cycles the validation
error was increasing. The picture gets even worse, if one looks at the fuH
crossvalidation results shown in Fig. 16.4. Obviously there are other splittings of
the training data, for which the error is about 50% of the data maximum, so that
the quality of the predictions depends strongly on the choice of the training data
and the predictions are not reliable. Accordingly, one has to conclude that for this
Neural Network generalization fails.
16.5
Conclusions
In this article we discussed the application of time series prediction by Neural
Networks to environmental data sets. Such data are typically of low quality as
compared to laboratory data for various reasons: boundary conditions for the
phenomenon to be studied are not controHable, taking clean data is much too
expensive or the phenomenon is so complex that the relevant variables that should
be measured are not known. The consequence is: one usually has to work with the
data one can get, and not the data one would like to have. The zooplankton
forecasts considered in the previous section illustrate this situation: the data are
