Chapter 14 . Time-Series Prediction of Chlorophyll a in Lakes
273
Table 14.2. Information regarding training of ANNs.
ANN feature
Value
Learning algorithm
SCG (Müller 1993).
Weight update mode Batch
Data conditioning
Input
- mean = 0; stdev = 1
Output - scaled to 0: 1
Transfer function
Sigmoid
Average training time< 10 seconds (DEC Alpha and
Intel P6 dass CPUs)
14.3.3
Control of Overfitting
It was elected to determine the effect of hidden layer configuration and training
time on the level of overfitting (variance) being expressed by the approximated
models. Hidden layer neurons each have the potential to define a single decision
boundary in the model with respect to the independent and dependent variables
(Wasserman 1989). The number of hidden layer neurons defines the number of
decision boundaries, and thus the complexity of the overall decision boundaries
represented by the model. Few decision boundaries will lead to a biased model
since it will be unable to describe all the information present. Many decision
boundaries lead to low bias, but potentially high variance.
Training time affects the extent to which the implicit complexity of an ANN
model is exploited. As training proceeds, the values of connection weights are
gradually increased and hidden layer neurons are progressively activated to form
decision boundaries in the model (Smith 1993). Stopping training early prevents
the potential decision boundaries from becoming fully expressed by constraining
connection weights to low values. "Early stopping" is thus referred to a nonconvergent means of biasing an ANN model (Finnoff et al. 1993).
Preliminary experimentation revealed that it was possible to achieve negligible
error on training sets (i.e no bias) for all available databases with 5 hidden layer
neurons using the learning techniques described in section 14.3.2. It was therefore
assumed that appropriate bias to minimise overall error could be achieved by
searching ANN architectures with 5 or less hidden neurons. A factorial
experimental design was implemented for the two input designs (same day and 30
day models) and all six databases utilising all combinations of hidden layer
configuration (0, 2, and 5 hidden nodes) and stopping errors of training (0, 0.5,
1.0, 1.5,2.0) giving a total of 180 treatments.
Précédent

- 291/410

Suivant