314
C.H. Reick· A. Grünewald . B. Page
Especially when low quality data are used for prediction or if the number of
simultaneous variables is high, it is often hard to judge, whether a particular
Neural Net is able to generalize. This situation is often encountered when working
with environmental data sets, as everyone knows, who tried to work e.g. with
biological time series (see e.g. Reick and Page (2000». The reasons for this low
quality are simple: First, environmental data are typically taken under nonlaboratory conditions, i.e. external disturbances cannot be controlled and the data
get noisy. Second, one has usually to measure what one can get, and not what one
would like to measure. So one cannot be sure, whether the data are representative
for some hidden deterministic dynamics. And finally, long measurement
campaigns are expensive so that environmental data sets are typically quite short
compared to their noise level. This problem is only insufficiently compensated by
measuring simultaneously several variables (the extra costs are typically low),
because hereby one can only improve the information on particular system states,
but cannot gain additional information on the diversity of system states; this could
only be obtained from sufficiently long time series. But this information on the
diversity of states is indispensable for predictions of high quality. Moreover, the
information of additional variables is often redundant and also noisy, to the
consequence that by using these data as additional inputs in Neural Nets their
performance can get even worse.
The prediction quality is usually measured by computing the average prediction
error for a number of prediction instants. But the question is, how far this
prediction error can be trusted, when a Neural Network is used to predict
unknown data. When working with Neural Networks one experiences that the
lower the data quality, the less reliable are the computed prediction errors. As
already discussed, this situation is especially encountered, when working with
environmental data so that here one should always carefully analyze their
reliability. This means one has to inquire the ability of a Neural Network to
generalize. How to do this by cross validation techniques will be discussed in
section 11. Unfortunately, crossvalidation is very laborous, because the same
Neural Network has to be trained over and over again with different parts of the
available data. The solution can only be a complete automatization of the training
process. Standard Neural Network software, like e.g. SNNS (Zell 1994), supports
mainly the visual supervision of the training at the computer monitor. But this is
much too laborous when performing crossvalidation studies. Alternatively, one
could use the programming interfaces, that are part of many Neural Network
products. But besides the uncomfortability of such a solution, there is a more
fundamental problem with automatization: Many training algorithms have been
developed in the past and many of them are available in Neural Network
packages. But when using them for automatized training the main problem is how
to stop the training, such that the network is neither under- and nor overadapted in
order to guarantee optimal generalization. For visual supervision at the screen,
there is a widely accepted stopping technique by Weigend et al. (1991). As a first
C.H. Reick· A. Grünewald . B. Page
Especially when low quality data are used for prediction or if the number of
simultaneous variables is high, it is often hard to judge, whether a particular
Neural Net is able to generalize. This situation is often encountered when working
with environmental data sets, as everyone knows, who tried to work e.g. with
biological time series (see e.g. Reick and Page (2000». The reasons for this low
quality are simple: First, environmental data are typically taken under nonlaboratory conditions, i.e. external disturbances cannot be controlled and the data
get noisy. Second, one has usually to measure what one can get, and not what one
would like to measure. So one cannot be sure, whether the data are representative
for some hidden deterministic dynamics. And finally, long measurement
campaigns are expensive so that environmental data sets are typically quite short
compared to their noise level. This problem is only insufficiently compensated by
measuring simultaneously several variables (the extra costs are typically low),
because hereby one can only improve the information on particular system states,
but cannot gain additional information on the diversity of system states; this could
only be obtained from sufficiently long time series. But this information on the
diversity of states is indispensable for predictions of high quality. Moreover, the
information of additional variables is often redundant and also noisy, to the
consequence that by using these data as additional inputs in Neural Nets their
performance can get even worse.
The prediction quality is usually measured by computing the average prediction
error for a number of prediction instants. But the question is, how far this
prediction error can be trusted, when a Neural Network is used to predict
unknown data. When working with Neural Networks one experiences that the
lower the data quality, the less reliable are the computed prediction errors. As
already discussed, this situation is especially encountered, when working with
environmental data so that here one should always carefully analyze their
reliability. This means one has to inquire the ability of a Neural Network to
generalize. How to do this by cross validation techniques will be discussed in
section 11. Unfortunately, crossvalidation is very laborous, because the same
Neural Network has to be trained over and over again with different parts of the
available data. The solution can only be a complete automatization of the training
process. Standard Neural Network software, like e.g. SNNS (Zell 1994), supports
mainly the visual supervision of the training at the computer monitor. But this is
much too laborous when performing crossvalidation studies. Alternatively, one
could use the programming interfaces, that are part of many Neural Network
products. But besides the uncomfortability of such a solution, there is a more
fundamental problem with automatization: Many training algorithms have been
developed in the past and many of them are available in Neural Network
packages. But when using them for automatized training the main problem is how
to stop the training, such that the network is neither under- and nor overadapted in
order to guarantee optimal generalization. For visual supervision at the screen,
there is a widely accepted stopping technique by Weigend et al. (1991). As a first
