Chapter 14 . Time-Series Prediction of Chlorophyll a in Lakes
267
• It potentially reduces modelling effort through rationalisation of monitoring,
database, and modelling approaches.
Maier et al. (1998) and Recknagel et al. (1998) introduced a time delay input
structure (Waibel 1989) utilising lagged inputs that are a given number of
timesteps in the past relative to the output. This idea exploits serial correlation in
the data, and accounts for time dynamic relationships between lake conditions,
growth and algal populations. Both of these studies report improved prediction
accuracy as a result. Additionally, Maier et al. (1998) demonstrated that such a
structure enables forecasts of algal abundance in rivers for up to four weeks into
the future. Therefore, it is proposed in the current work to test the hypothesis that
utilisation of lag inputs will confer similar advantages to the generic ANN
structure applied to a wider range of case studies.
14.2.2
Control of Overfitting
One advantage of ANNs, GAs and other so-called non-parametric modelling
approaches is that they are characterised by asymptotic convergence on the model
underlying a given population of data with increasing training sampie size (c.f. the
property of consistency) (Geman et al. 1992). However, the amount of data
required to achieve convergence may be i) very difficult to determine analytically
(Moody 1991) and ii) frequently impractically large for many applications
(Geman et al. 1992). The primary consequence of non-convergence through
inadequate training sampie representation is high error due to model variance
arising from the sampling variability of training sets. This high variance
component of error is generally referred to as overfltting (Moody 1991).
Control of overfitting behaviour (i.e variance) is usually achieved by the
employment of an arbitrary bias. Since introducing a bias in the form of a model
is generally considered to be defeating the purpose of the exercise (since we are
trying to learn something new from data), the usual approach is to employ some
penalisation of the ANN approximation process such that it is constrained to
smoother, less complex decision boundaries. Penalisation techniques include:
1. Limitation of hidden layer size (either a-priori, or on-line by pruning
connections) thus lowering the Vapnik-Chervonenkis (VC) dimension of the
ANN (Abu-Mustafa 1989).
2. Early stopping of training (Finnoff 1993).
3. Introduction of a regularisation term by weight decay.
4. Use of a training algorithm predisposed towards good generalisation such as
backpropagation (Lawrence and Giles 2000).
The trade-off involved with penalisation of the ANN model approximation is of
course increased model error due to bias. Thus a key problem facing ANN users
in practical applications is determination of how much the ANN should be biased
to minimise the total (bias + variance) error (see Figure 14.1). In general,
practitioners either use some heuristic, or perform a trial and error configuration of
the relevant meta-parameters. Both of these approaches are problematic:
267
• It potentially reduces modelling effort through rationalisation of monitoring,
database, and modelling approaches.
Maier et al. (1998) and Recknagel et al. (1998) introduced a time delay input
structure (Waibel 1989) utilising lagged inputs that are a given number of
timesteps in the past relative to the output. This idea exploits serial correlation in
the data, and accounts for time dynamic relationships between lake conditions,
growth and algal populations. Both of these studies report improved prediction
accuracy as a result. Additionally, Maier et al. (1998) demonstrated that such a
structure enables forecasts of algal abundance in rivers for up to four weeks into
the future. Therefore, it is proposed in the current work to test the hypothesis that
utilisation of lag inputs will confer similar advantages to the generic ANN
structure applied to a wider range of case studies.
14.2.2
Control of Overfitting
One advantage of ANNs, GAs and other so-called non-parametric modelling
approaches is that they are characterised by asymptotic convergence on the model
underlying a given population of data with increasing training sampie size (c.f. the
property of consistency) (Geman et al. 1992). However, the amount of data
required to achieve convergence may be i) very difficult to determine analytically
(Moody 1991) and ii) frequently impractically large for many applications
(Geman et al. 1992). The primary consequence of non-convergence through
inadequate training sampie representation is high error due to model variance
arising from the sampling variability of training sets. This high variance
component of error is generally referred to as overfltting (Moody 1991).
Control of overfitting behaviour (i.e variance) is usually achieved by the
employment of an arbitrary bias. Since introducing a bias in the form of a model
is generally considered to be defeating the purpose of the exercise (since we are
trying to learn something new from data), the usual approach is to employ some
penalisation of the ANN approximation process such that it is constrained to
smoother, less complex decision boundaries. Penalisation techniques include:
1. Limitation of hidden layer size (either a-priori, or on-line by pruning
connections) thus lowering the Vapnik-Chervonenkis (VC) dimension of the
ANN (Abu-Mustafa 1989).
2. Early stopping of training (Finnoff 1993).
3. Introduction of a regularisation term by weight decay.
4. Use of a training algorithm predisposed towards good generalisation such as
backpropagation (Lawrence and Giles 2000).
The trade-off involved with penalisation of the ANN model approximation is of
course increased model error due to bias. Thus a key problem facing ANN users
in practical applications is determination of how much the ANN should be biased
to minimise the total (bias + variance) error (see Figure 14.1). In general,
practitioners either use some heuristic, or perform a trial and error configuration of
the relevant meta-parameters. Both of these approaches are problematic:
