Chapter 16 . Time-Series Prediction of Marine Zooplankton
321
("MAXCYC'). So far this algorithm describes simply a search for a best network
that is stopped after a fixed number of training cycles and it is completely
irrelevant whether the validation error has aminimum, is ever increasing or shows
a noisy behaviour with many minima. Insofar it has nothing to do with early
stopping. Nevertheless, if MAXCYC is set to a very large value so that one can be
sure that the global minimum of the validation error is reached (if it can be
reached), this algorithm gives the same result as early stopping. But numerically it
is not desireable to train the network in every case by the maximum number of
cycles. Instead, one would like to stop the training already when the optimal
network has been found. This cannot be fully achieved, but at least partially by
including the so far ignored parts into the algorithm.
To explain their function, we first suppose TOL to be zero. When starting the
algorithm the delay counter "D" is set to zero. If during training the validation
error decreases ("bestErr> err") the delay counter stays zero ("D=O"). But if the
error gets worse, i.e. if the validation error curve undergoes aminimum, the delay
counter starts counting ("D=D+ 1"). If the validation error continues to increase,
the delay counter will reach its maximum value DEIAY and the training is
stopped. So this is not really early stopping, but a delayed early stopping. The
reason to perform some additional training cycles is to check whether the detected
minimum is only a small local minimum or can be considered a global minimum.
To understand how the algorithm detects this situation, assume that the delay
counter already started counting, i.e. the network error already increased for D
training cycles, and that in the present training cycle a validation error is found
that is smaller than a11 previous validation errors ("err < bestErr"). Fig. 16.1
shows that in this case the delay counter is once more set to zero ("D=O").
Thereby the delay counter has to start once more from the beginning and the
training will last at least DEIAY additional training cycles (if not MAXCYC is
reached before). Thereby the local minimum has been eseaped. This means that
for TOL equal to zero, the training stops only, if a detected minimum remains the
absolute minimum for the next DEIAY training cycles. If not, the next minimum
will be checked. In this way training eseapes, as desired, loeal minima.
Nevertheless, there is a problem: What happens, if the validation error is ever
decreasing, or, decreasing on the average with local minima separated not farther
than DEIAY training cycles apart? In that case D would always be reset to zero,
before DEIAY has been reached. Accordingly, the algorithm continues training
for the fu11 MAXCYC training cycles. Such a behaviour is not always desireable,
namely, when the error improves only insignificantly, so that even by longer
training a substantia11y better network cannot be expected. To stop training in such
cases before MAXCYC has been reached, the additional parameter T OL has been
introduced. Once more the situation of a monotoneously decreasing validation
error is considered, but now with TOL larger than zero. In this case the delay
counter is reset to zero only if the validation error improves by an amont larger
than TOL, i.e. in the case "err < bestErr - TOL". Therefore the training is stopped,
321
("MAXCYC'). So far this algorithm describes simply a search for a best network
that is stopped after a fixed number of training cycles and it is completely
irrelevant whether the validation error has aminimum, is ever increasing or shows
a noisy behaviour with many minima. Insofar it has nothing to do with early
stopping. Nevertheless, if MAXCYC is set to a very large value so that one can be
sure that the global minimum of the validation error is reached (if it can be
reached), this algorithm gives the same result as early stopping. But numerically it
is not desireable to train the network in every case by the maximum number of
cycles. Instead, one would like to stop the training already when the optimal
network has been found. This cannot be fully achieved, but at least partially by
including the so far ignored parts into the algorithm.
To explain their function, we first suppose TOL to be zero. When starting the
algorithm the delay counter "D" is set to zero. If during training the validation
error decreases ("bestErr> err") the delay counter stays zero ("D=O"). But if the
error gets worse, i.e. if the validation error curve undergoes aminimum, the delay
counter starts counting ("D=D+ 1"). If the validation error continues to increase,
the delay counter will reach its maximum value DEIAY and the training is
stopped. So this is not really early stopping, but a delayed early stopping. The
reason to perform some additional training cycles is to check whether the detected
minimum is only a small local minimum or can be considered a global minimum.
To understand how the algorithm detects this situation, assume that the delay
counter already started counting, i.e. the network error already increased for D
training cycles, and that in the present training cycle a validation error is found
that is smaller than a11 previous validation errors ("err < bestErr"). Fig. 16.1
shows that in this case the delay counter is once more set to zero ("D=O").
Thereby the delay counter has to start once more from the beginning and the
training will last at least DEIAY additional training cycles (if not MAXCYC is
reached before). Thereby the local minimum has been eseaped. This means that
for TOL equal to zero, the training stops only, if a detected minimum remains the
absolute minimum for the next DEIAY training cycles. If not, the next minimum
will be checked. In this way training eseapes, as desired, loeal minima.
Nevertheless, there is a problem: What happens, if the validation error is ever
decreasing, or, decreasing on the average with local minima separated not farther
than DEIAY training cycles apart? In that case D would always be reset to zero,
before DEIAY has been reached. Accordingly, the algorithm continues training
for the fu11 MAXCYC training cycles. Such a behaviour is not always desireable,
namely, when the error improves only insignificantly, so that even by longer
training a substantia11y better network cannot be expected. To stop training in such
cases before MAXCYC has been reached, the additional parameter T OL has been
introduced. Once more the situation of a monotoneously decreasing validation
error is considered, but now with TOL larger than zero. In this case the delay
counter is reset to zero only if the validation error improves by an amont larger
than TOL, i.e. in the case "err < bestErr - TOL". Therefore the training is stopped,
