216
G.J. Bowden . G.C. Dandy· H.R. Maier
estimating the unknown model parameters. Data driven approaches are usually
assumed to be able to determine which model inputs are critical. However, as
pointed out by Maier and Dandy (2000b), presenting a large number of inputs to
the ANN and relying on the network to determine the significant inputs, usually
increases the network size. This results in certain disadvantages, such as an
increase in the amount of data required to estimate the connection weights
properly and a reduction in processing speed (Lachtermacher and Fuller 1994).
The input selection problem can be formulated as having a set of input
variables to an ANN, and an output value which can be used to evaluate the fitness
or merit of the network using those variables. For example, for the problem of
forecasting cyanobacteria, the inputs to the ANN are causal variables such as
turbulence (fIow), water temperature, turbidity, colour and nitrogen and
phosphorus concentrations. From these variables, a subset of inputs must be
selected that yield the network of highest fitness. This subset of inputs forms an
n-dimensional input vector X", which can be used to forecast the concentration of
cyanobacteria Y. The aim of the ANN is to produce a generalised relationship of
the form
Y=j(X")
(11.1)
It is a generalised relationship because the functional form of j(.) is not
revealed explicitly but rather, is represented by the ANN model's structure and
parameters. The network's fitness can be determined using an appropriate
measure e.g. smallest root mean square prediction erroT. In complex applications,
the number of input variables can be quite large and this problem is further
exacerbated in time series studies, where appropriate lags must also be chosen.
Therefore, analytical techniques for determining the optimal subset of inputs
present the modeller with a distinct advantage.
Maier and Dandy (2000b) reviewed 43 papers on the application of ANNs for
hydrological modelling, and found that in many cases the lack of a methodology
to determine the input variables raised doubt about the optimality of the inputs
obtained. In some instances, inputs were chosen arbitrarily. In other cases, a
priori knowledge was used and when different methods were employed, such as
trial-and-error, often the validation data were used as part of the training process.
Intuitively, the preferred approach for determining appropriate inputs and lags of
inputs, involves a combination of apriori knowledge and analytical approaches
(Maier and Dandy 1997; Fernando and Jayawardena 1998; Maier et al. 1998).
There are two broad stages in input determination. Firstly, unsupervised input
preprocessing (i.e. discarding redundant inputs) and secondly, supervised input
selection (i.e. using the ANN's output in an analytical procedure to determine the
significant input variables). In unsupervised input preprocessing, the original set
of input variables is processed to produce a subset of inputs containing as much
information as possible from the original set. This sub set of inputs can then be
used in a supervised input selection
process to determine which combinations of these inputs result in the network
of highest fitness.
G.J. Bowden . G.C. Dandy· H.R. Maier
estimating the unknown model parameters. Data driven approaches are usually
assumed to be able to determine which model inputs are critical. However, as
pointed out by Maier and Dandy (2000b), presenting a large number of inputs to
the ANN and relying on the network to determine the significant inputs, usually
increases the network size. This results in certain disadvantages, such as an
increase in the amount of data required to estimate the connection weights
properly and a reduction in processing speed (Lachtermacher and Fuller 1994).
The input selection problem can be formulated as having a set of input
variables to an ANN, and an output value which can be used to evaluate the fitness
or merit of the network using those variables. For example, for the problem of
forecasting cyanobacteria, the inputs to the ANN are causal variables such as
turbulence (fIow), water temperature, turbidity, colour and nitrogen and
phosphorus concentrations. From these variables, a subset of inputs must be
selected that yield the network of highest fitness. This subset of inputs forms an
n-dimensional input vector X", which can be used to forecast the concentration of
cyanobacteria Y. The aim of the ANN is to produce a generalised relationship of
the form
Y=j(X")
(11.1)
It is a generalised relationship because the functional form of j(.) is not
revealed explicitly but rather, is represented by the ANN model's structure and
parameters. The network's fitness can be determined using an appropriate
measure e.g. smallest root mean square prediction erroT. In complex applications,
the number of input variables can be quite large and this problem is further
exacerbated in time series studies, where appropriate lags must also be chosen.
Therefore, analytical techniques for determining the optimal subset of inputs
present the modeller with a distinct advantage.
Maier and Dandy (2000b) reviewed 43 papers on the application of ANNs for
hydrological modelling, and found that in many cases the lack of a methodology
to determine the input variables raised doubt about the optimality of the inputs
obtained. In some instances, inputs were chosen arbitrarily. In other cases, a
priori knowledge was used and when different methods were employed, such as
trial-and-error, often the validation data were used as part of the training process.
Intuitively, the preferred approach for determining appropriate inputs and lags of
inputs, involves a combination of apriori knowledge and analytical approaches
(Maier and Dandy 1997; Fernando and Jayawardena 1998; Maier et al. 1998).
There are two broad stages in input determination. Firstly, unsupervised input
preprocessing (i.e. discarding redundant inputs) and secondly, supervised input
selection (i.e. using the ANN's output in an analytical procedure to determine the
significant input variables). In unsupervised input preprocessing, the original set
of input variables is processed to produce a subset of inputs containing as much
information as possible from the original set. This sub set of inputs can then be
used in a supervised input selection
process to determine which combinations of these inputs result in the network
of highest fitness.
