Chapter 16 . Time-Series Prediction of Marine Zooplankton
323
moderate prediction quality is achievable. For a given abundance x we define this
magnitude m by
m = [ogu!,x+ 1)
(16.1).
(The addition of "1" in this equation assures that abundance zero is mapped to
magnitude zero; for large abundance the magnitude differs only negligibly from
its decadic logarithm.) Accordingly, in the following we will consider only
predictions of the magnitude of abundances and not of the abundances itself.
TechnicaIly this was achieved by transforming all plankton data from abundances
to magnitudes before averaging. This has also been done with the phytoplankton
data.
It is generally expected that there is a complex network of interactions between
the several types of plankton organisms ("food web"). Accordingly we tried to
perform predictions by using simultaneously the time series of several taxa as
input of the Neural Networks. In view of the large amount of possible
combinations of taxa such studies could only be performed by massive
automatization. Our aim here is not to present the results from this study (see
Grünewald 2000). Instead we want to demonstrate the application of
crossvalidation and early stopping to a particular prediction problem. The example
we consider is the prediction of BarnacIe larvae (Cirripedia nauplius) with a
16x5x2xl feedforward Neural Network. As input we use the data from the last
eight weeks of Cirripedia nauplius itself and also the last eight weeks of our
diatom data. This combination was chosen because Cirripedia nauplius at least
partially preys on diatoms. As learning algorithm we used "Resilient
Backpropagation" (Riedmiller and Braun 1993), which is usually faster than the
cIassical backpropagation algorithm.
From the 20 years of data we used the first 16 years as in-sample data and the
last four years as out-of-sample data. For training the in-sample data were
partititioned in 30 different ways into training and validation data (see Fig. 16.2);
we used 12 years for training and 4 years for validation. The parameters used for
the early stopping algorithm were MAXCYC = 1000, MAXDEL = 100 and TOL =
10
04 • Fig. 16.3 shows the predictions for the out-of-sample data obtained by the
Neural Network that was trained with the data splitting no. 16 of Fig. 16.2. The
predictions reproduce quite well the data. This is substantiated by a value of 0.91
for the correlation between the predictions and the data. Accordingly, the error,
also shown in Fig. 16.3, is on the average significantly smaIler than the data,
although not small. At wintertime, where the error is of the same order as the data,
better prediction results cannot be expected, because there are often less than 15
individuals per m
3 present, so that already the statistical uncertainty from sampling
is of the order of the abundanceso Overall, this prediction looks quite weIl.
323
moderate prediction quality is achievable. For a given abundance x we define this
magnitude m by
m = [ogu!,x+ 1)
(16.1).
(The addition of "1" in this equation assures that abundance zero is mapped to
magnitude zero; for large abundance the magnitude differs only negligibly from
its decadic logarithm.) Accordingly, in the following we will consider only
predictions of the magnitude of abundances and not of the abundances itself.
TechnicaIly this was achieved by transforming all plankton data from abundances
to magnitudes before averaging. This has also been done with the phytoplankton
data.
It is generally expected that there is a complex network of interactions between
the several types of plankton organisms ("food web"). Accordingly we tried to
perform predictions by using simultaneously the time series of several taxa as
input of the Neural Networks. In view of the large amount of possible
combinations of taxa such studies could only be performed by massive
automatization. Our aim here is not to present the results from this study (see
Grünewald 2000). Instead we want to demonstrate the application of
crossvalidation and early stopping to a particular prediction problem. The example
we consider is the prediction of BarnacIe larvae (Cirripedia nauplius) with a
16x5x2xl feedforward Neural Network. As input we use the data from the last
eight weeks of Cirripedia nauplius itself and also the last eight weeks of our
diatom data. This combination was chosen because Cirripedia nauplius at least
partially preys on diatoms. As learning algorithm we used "Resilient
Backpropagation" (Riedmiller and Braun 1993), which is usually faster than the
cIassical backpropagation algorithm.
From the 20 years of data we used the first 16 years as in-sample data and the
last four years as out-of-sample data. For training the in-sample data were
partititioned in 30 different ways into training and validation data (see Fig. 16.2);
we used 12 years for training and 4 years for validation. The parameters used for
the early stopping algorithm were MAXCYC = 1000, MAXDEL = 100 and TOL =
10
04 • Fig. 16.3 shows the predictions for the out-of-sample data obtained by the
Neural Network that was trained with the data splitting no. 16 of Fig. 16.2. The
predictions reproduce quite well the data. This is substantiated by a value of 0.91
for the correlation between the predictions and the data. Accordingly, the error,
also shown in Fig. 16.3, is on the average significantly smaIler than the data,
although not small. At wintertime, where the error is of the same order as the data,
better prediction results cannot be expected, because there are often less than 15
individuals per m
3 present, so that already the statistical uncertainty from sampling
is of the order of the abundanceso Overall, this prediction looks quite weIl.
