Chapter 7 . Stream Ecosystem Analysis
123
their origin either in an unreliable model or in biased, unrepresentative data.
Similarly, a smaH error may indicate a good model or, regarding outliers, a good
ans wer just by chance. Therefore, the error measure can only help during the
training phase. When the network deals with new input data, no output
observations are available and hence the error can only be estimated on the basis
of the local variability of output values corresponding to training inputs dose to
the actual input. Furthermore, the local reliability of the model depends on the
amount of training data similar to the actual input vector and the distance to the
nearest training vector.
We extended the combination of SOMs and RBF Networks to estimate the
reliability of the network outputs and to get a better insight into the internat
network activities. SOM-training optimised the centres of radial basis functions of
an RBF Neural Network (Bishop 1995) and provided visualisation of the RBFlayer's activation patterns.
In analogy to the above studies, we predicted oxygen concentration based on
macro-invertebrate species abundance (Werner and Obach 2001). Six test sites of
51 were randomly chosen. Fourteen predictors were selected by GRNNs and a
genetic algorithm, which overcame the stepwise method.
The activations of the RBF neurons can be displayed for any input pattern on
the SOM, wh ich has been used previously for optimising the centres of radial
basis functions. Thus it is possible to visualize and decide whether a test site is
weH known to the network or whether an extrapolation must be assumed (Werner
and Obach 2001).
The variability of the output variable in dasses determined by the SOM, can be
displayed by box-and-whisker plots (Obach et al. 2001).
7.6
Conclusions
Artificial Neural Networks proved to be suitable tools to model non-linear
interrelations in basic and applied ecology of running waters. Although they are
able to describe correlations between multi-dimensional variables, causality
detection is not possible. The major goal of model building is generalisation. Data
sampies must comprise as many information as possible and be representative for
training and testing ANNs. For smaH data sets the performance of fast leaming
networks (LNN, GRNN, RBF) was estimated by cross-validation. Large amounts
of data require pre-processing. Dimension reduction was performed with
correlation analysis, multiple linear regression, principal component analysis,
bottle-neck networks, sensitivity analysis, genetic algorithms and stepwise
methods, depending on the applied ANN type and the linearity of the described
correlation. Linear networks are simple in use, training and interpretation. They
provide a benchmark against which the quality of other models like MLPs,
GRNNs, RBFNs, MFMs is compared. The application of different network types
on the same problem is recommended.
Précédent

- 145/410

Suivant