110
I.M. Schleiter . M. Obach . R. Wagner· H. Werner· H.-H. Schmidt
. D. Borchardt
variables for modelling, e.g. indicator species prediction of ecological properties
of lotic ecosystems prediction of the assemblage of benthic communities in
disturbed and undisturbed streams generalisation of interdependencies.
In an interdisciplinary research project of mathematicians, computer scientists,
ecologists, and engineers, the suitability of various types of ANNs was tested.
They were used to model temporal dynamics of water quality based on weather,
urban storm-water run-off and waste-water effluents, bioindication of lotic
ecosystem properties using benthic macroinvertebrates, and long-term population
dynamics of aquatic insects depending on environmental and ecological variables.
7.2
Materials and Methods
7.2.1
Data Base
Our research was based on two data sets from running waters in Hesse (Central
Germany):
Nine streams with different amounts of organic pollution (Schleiter et al. 1999;
2001) (248 macro-zoobenthos taxa and physical, chemical, and hydromorphological variables)
A thirty years data set of environmental variables (precipitation, discharge,
water temperature) and aquatic insects (Ephemeroptera, Plecoptera, Trichoptera)
of an almost pristine stream (Limnological River Station Schlitz; Obach et al.
2001).
7.2.2
Data Pre-Processing
Pre-processing includes all data alterations before the applications of ANNs. The
first step is a test of completeness, e.g. determination of an adequate method to
handle missing values, and plausibility. Outliers can be detected and mapped onto
the borders of a reliable range of values (truncation) or use non-linear functions,
e.g. logarithmic or sigmoidal.
Variables can be normalized in order to avoid an undesireably high influence of
large absolute values. We usually mapped the values linearly onto the interval
[0,1]; occasionally standardisation is more advisable.
The data set was usually divided into training (adaption of net parameters),
verification (selection of an adequate model) and test data (to estimate
generalisation ability).
Selection of the most relevant variables by regression, correlation analyses or
based on expert knowledge or combination of variables by Principal Component
I.M. Schleiter . M. Obach . R. Wagner· H. Werner· H.-H. Schmidt
. D. Borchardt
variables for modelling, e.g. indicator species prediction of ecological properties
of lotic ecosystems prediction of the assemblage of benthic communities in
disturbed and undisturbed streams generalisation of interdependencies.
In an interdisciplinary research project of mathematicians, computer scientists,
ecologists, and engineers, the suitability of various types of ANNs was tested.
They were used to model temporal dynamics of water quality based on weather,
urban storm-water run-off and waste-water effluents, bioindication of lotic
ecosystem properties using benthic macroinvertebrates, and long-term population
dynamics of aquatic insects depending on environmental and ecological variables.
7.2
Materials and Methods
7.2.1
Data Base
Our research was based on two data sets from running waters in Hesse (Central
Germany):
Nine streams with different amounts of organic pollution (Schleiter et al. 1999;
2001) (248 macro-zoobenthos taxa and physical, chemical, and hydromorphological variables)
A thirty years data set of environmental variables (precipitation, discharge,
water temperature) and aquatic insects (Ephemeroptera, Plecoptera, Trichoptera)
of an almost pristine stream (Limnological River Station Schlitz; Obach et al.
2001).
7.2.2
Data Pre-Processing
Pre-processing includes all data alterations before the applications of ANNs. The
first step is a test of completeness, e.g. determination of an adequate method to
handle missing values, and plausibility. Outliers can be detected and mapped onto
the borders of a reliable range of values (truncation) or use non-linear functions,
e.g. logarithmic or sigmoidal.
Variables can be normalized in order to avoid an undesireably high influence of
large absolute values. We usually mapped the values linearly onto the interval
[0,1]; occasionally standardisation is more advisable.
The data set was usually divided into training (adaption of net parameters),
verification (selection of an adequate model) and test data (to estimate
generalisation ability).
Selection of the most relevant variables by regression, correlation analyses or
based on expert knowledge or combination of variables by Principal Component
