118
I.M. Schleiter . M. Obach . R. Wagner' H. Wemer' H.-H. Schmidt
. D. Borchardt
The determination coefficient B (on the 20% test data) was surprisingly high. It
was >0.9 for five species, >0.7 for 9 species and >0.5 for 15 species. Only for P.
intricata and P. auberti was B just below 0.5. The quality of the prognoses varied
among species and among the different ANNs. This was either due to the variable
abundances in the test data, or was related to the ecological plasticity of the
populations. As expected, linear regression models had the lowest B values in
almost all species. In A. jimbriata, C. villosa, T. rostocki, P. auberti and P.
intricata linear regression and ANN models with pre-selection based on regression
were of similar quality. Pre-selection of the five best variables by regression
analysis (8 times) or by sensitivity analysis (6 times) improved prognoses. In three
of seventeen species the reduction to the best five predictors was not accompanied
by an increase of B. Only in the model for L. nigra did the use of all (51)
predictors drastically increase B (by 0.29 or even 0.62) compared to models with
the five most important variables. This may be due to the species life-cycle
attributes (species traits). Larvae live on and in the relatively unstable sandy
sediments and thus habitat and specimens are susceptible to almost every change
in discharge throughout the year.
In summary, for most species models with dimension reduction by regression
or sensitivity analysis produced models of similar quality (i.e. difference of
BeO.05). Pronounced differences were found for L. prima and P. auberti (preselection by sensitivity analysis), or l. goertzi, P. intricata, S. torrentium and T.
rostocki (pre-selection by stepwise linear regression).
Reliability of all models was tested on the background of ANN computation
and ecological knowledge. The results indicated that enhanced ecological
flexibility of populations (risk spreading), low temporal resolution of the data,
data scaling method, or different occurrences in leaming and testing data resulted
in a low model quality.
Scaling during pre-processing is one crucial step in exploring ecological data,
and subsequent modelling. We transformed values linearly, sigmoidally,
logarithmically and exponentially. Logarithmic scaling was optimal for discharge,
to smooth extreme or rare events (floods). Prognoses for the abundance of Baetis
vernus (and B. rhodani) after logarithmic scaling of all predictors resulted in a
B=O.77 (compare Table 7.1 with linear scaled variables). However, after
transforming the predicted values into original units, B was 0.63 or lower.
Therefore, it appears that non-linear scaling did not improve any model.
Even though four different models with high determination coefficients were
developed for four individual study sites (and six species), generalisation ability of
the models was not as expected. Extrapolation on neighbouring sites at 600 m
distance upstream or downstream was restricted.
Extreme values or low data density occurred in the different training data sets.
If data points are dissimilar to the trained model, they should be attributed as
novelty. To recognize those cases in general is part of the future work (novelty
detection). The probability that many zero-predictions (i.e. no abundance in a
particular month) had artificially increased the model quality and led to different
approaches (Obach et al. 2001). In addition, recurrent ANNs of the Jordan and
Elman types were used without significant success.
Précédent

- 140/410

Suivant