Chapter 16 . Time-Series Prediction of Marine Zooplankton
315
step to automatization we show in section III how this technique can be cast into
an algorithm. Finally we show in section IV how crossvalidation and automatized
training can be applied to an environmental data set, namely to plankton time
series from the North Sea. In contrast to most other prediction studies, we will not
show how well Neural Nets predict these time series, but instead show how the
failure of their ability to generalize can be substantiated.
16.2
Generalization
To find a Neural Network, by which a particular time series can be correctly
predicted, many prediction experiments have to be performed. In these
experiments one varies the type of input data, modifies the preprocessing of the
data and changes the internal structure of the Neural Networks. The success of a
particular prediction experiment is usually measured by the prediction error that is
obtained when trying to predict data that were not used during training. The
Neural Network with the least prediction error will then be chosen to perform
actual predictions.
When working with data of high quality, this procedure often works quite weil.
But when working with poor data the prediction quality is low, and this leads to
two problems. First, the experimental effort to find a Neural Network with an
acceptable prediction error increases. Whereas this is mainly a practical problem,
the second is of more fundamental nature: The reliability of the predictions gets
questionable. In the case of a Neural Network with a small prediction error (small
in relation to a characteristic scale of the data; the error computed for available
data, not in actual predictions) a doubling or even tripling of the error would still
be a small error. Therefore, a small prediction error is a good indication that for
actual predictions the error will also be small. For low quality data the situation is
different. If one finds only networks with prediction errors that cannot be judged
small, e.g. 20-30% relative error, then a doubling or tripling of the error in actual
predictions will no more be acceptable. Therefore, in this case, the reliability of
the predictions gets problematic.
In our opinion, this problem can be discussed in the context of the more general
problem of generalization. The term "generalization" is usually bound to a small
output error (here: prediction error) for non-training data. It is dear that a small
prediction error is an indication of a successful generalization. But what about
generalization, if the error is not small? Even in that case a Neural Network may
correctly reproduce essential features of a time series - although with a
significant error. Obviously in such a situation other aspects than the prediction
error get relevant for the question of generalization, as e.g. the reliability of the
prediction error, as discussed above. Therefore, in the following, we will discuss
how a generalization success or failure can be detected for low quality predictions.
315
step to automatization we show in section III how this technique can be cast into
an algorithm. Finally we show in section IV how crossvalidation and automatized
training can be applied to an environmental data set, namely to plankton time
series from the North Sea. In contrast to most other prediction studies, we will not
show how well Neural Nets predict these time series, but instead show how the
failure of their ability to generalize can be substantiated.
16.2
Generalization
To find a Neural Network, by which a particular time series can be correctly
predicted, many prediction experiments have to be performed. In these
experiments one varies the type of input data, modifies the preprocessing of the
data and changes the internal structure of the Neural Networks. The success of a
particular prediction experiment is usually measured by the prediction error that is
obtained when trying to predict data that were not used during training. The
Neural Network with the least prediction error will then be chosen to perform
actual predictions.
When working with data of high quality, this procedure often works quite weil.
But when working with poor data the prediction quality is low, and this leads to
two problems. First, the experimental effort to find a Neural Network with an
acceptable prediction error increases. Whereas this is mainly a practical problem,
the second is of more fundamental nature: The reliability of the predictions gets
questionable. In the case of a Neural Network with a small prediction error (small
in relation to a characteristic scale of the data; the error computed for available
data, not in actual predictions) a doubling or even tripling of the error would still
be a small error. Therefore, a small prediction error is a good indication that for
actual predictions the error will also be small. For low quality data the situation is
different. If one finds only networks with prediction errors that cannot be judged
small, e.g. 20-30% relative error, then a doubling or tripling of the error in actual
predictions will no more be acceptable. Therefore, in this case, the reliability of
the predictions gets problematic.
In our opinion, this problem can be discussed in the context of the more general
problem of generalization. The term "generalization" is usually bound to a small
output error (here: prediction error) for non-training data. It is dear that a small
prediction error is an indication of a successful generalization. But what about
generalization, if the error is not small? Even in that case a Neural Network may
correctly reproduce essential features of a time series - although with a
significant error. Obviously in such a situation other aspects than the prediction
error get relevant for the question of generalization, as e.g. the reliability of the
prediction error, as discussed above. Therefore, in the following, we will discuss
how a generalization success or failure can be detected for low quality predictions.
