Chapter 6 . Stream Assessment
95
are created based on that threshold. For the subsets of training examples in each
branch, the tree construction algorithm is called recursively. Tree construction
stops when all examples in anode are of the same dass (or if some other stopping
criterion is satisfied). Such nodes are called leaves and are labelled with the
corresponding values of the dass.
A number of systems exist for inducing dassification trees from examples, e.g.,
CART (Breiman et al. 1984), ASSISTANT (Cestnik et al. 1987), and C4.5
(Quinlan 1993). Of these, C4.5 is one of the most well known and widely-used
decision tree induction systems. J48 (Witten and Frank 1999) is a Java reimplementation of C4.5. It is apart of the machine learning package WEKA,
which also includes some of the latest developments in machine learning. This J48
was also used for inducing dassification trees and prediction models of
macroinvertebrate taxa of the Zwalm river basin. Standard settings were used. The
model validation was based on splitting the dataset in a training and validation set
of respectively 40 and 20 instances. Also ten fold cross validation on the whole
dataset (60 instances) was applied in specific cases.
6.2.4
Artificial Neural Networks
Artificial neural networks (ANNs) are mathematical models based on the transfer
of information through a network of functional units, called neurons. Given a
number of input values, entered at the basis of the network, it generates one or
more outputs. ANNs are currently recognized as an alternative for multivariate
statistics to predict aquatic communities (Gabriels et al. 2000). Recently, several
studies have been published, concerning the application of neural networks for
relating freshwater macroinvertebrates with their abiotic environment (e.g. Walley
and Fontama 1998; Schleiter et al. 1999; Gabriels et al. 2000). The neural network
was implemented with the neural network extension of the software package
MATLAB 5.3 for MS Windows™ (Demuth and Beale 1998). The model
validation was based on splitting the dataset in a training and validation set of
respectively 40 and 20 instances. Also 10 fold cross validation was sometimes
applied, as described by Witten and Frank (2000). Several optimisation studies
were carried out to select the best model configuration (Dedecker 2001). The best
neural network consisted of one hidden layer and ten neurons, with 'tansig' and
'logsig' transfer functions and 'gradient descending with momentum and adaptive
leaming rate backpropagation' as training algorithm (Demuth and Beale 1998). A
scheme of this neural network is shown in Fig. 6.3. During the research, special
interest was paid to the influence of the frequency of occurrence of the taxa on the
prediction reliability of the developed models. Sensitivity analyses were used to
get insight in the applied 'concepts' of these black box models.
95
are created based on that threshold. For the subsets of training examples in each
branch, the tree construction algorithm is called recursively. Tree construction
stops when all examples in anode are of the same dass (or if some other stopping
criterion is satisfied). Such nodes are called leaves and are labelled with the
corresponding values of the dass.
A number of systems exist for inducing dassification trees from examples, e.g.,
CART (Breiman et al. 1984), ASSISTANT (Cestnik et al. 1987), and C4.5
(Quinlan 1993). Of these, C4.5 is one of the most well known and widely-used
decision tree induction systems. J48 (Witten and Frank 1999) is a Java reimplementation of C4.5. It is apart of the machine learning package WEKA,
which also includes some of the latest developments in machine learning. This J48
was also used for inducing dassification trees and prediction models of
macroinvertebrate taxa of the Zwalm river basin. Standard settings were used. The
model validation was based on splitting the dataset in a training and validation set
of respectively 40 and 20 instances. Also ten fold cross validation on the whole
dataset (60 instances) was applied in specific cases.
6.2.4
Artificial Neural Networks
Artificial neural networks (ANNs) are mathematical models based on the transfer
of information through a network of functional units, called neurons. Given a
number of input values, entered at the basis of the network, it generates one or
more outputs. ANNs are currently recognized as an alternative for multivariate
statistics to predict aquatic communities (Gabriels et al. 2000). Recently, several
studies have been published, concerning the application of neural networks for
relating freshwater macroinvertebrates with their abiotic environment (e.g. Walley
and Fontama 1998; Schleiter et al. 1999; Gabriels et al. 2000). The neural network
was implemented with the neural network extension of the software package
MATLAB 5.3 for MS Windows™ (Demuth and Beale 1998). The model
validation was based on splitting the dataset in a training and validation set of
respectively 40 and 20 instances. Also 10 fold cross validation was sometimes
applied, as described by Witten and Frank (2000). Several optimisation studies
were carried out to select the best model configuration (Dedecker 2001). The best
neural network consisted of one hidden layer and ten neurons, with 'tansig' and
'logsig' transfer functions and 'gradient descending with momentum and adaptive
leaming rate backpropagation' as training algorithm (Demuth and Beale 1998). A
scheme of this neural network is shown in Fig. 6.3. During the research, special
interest was paid to the influence of the frequency of occurrence of the taxa on the
prediction reliability of the developed models. Sensitivity analyses were used to
get insight in the applied 'concepts' of these black box models.
