Chapter 15 . Evolved Predictive Rules for Aigal Dynamics
293
potentially toxic blue-green algae formed outbreaks in the lake during summer.
While Microcystis spp bloomed year by year until 1986, filamentous blue-green
algae such as Oscillatoria and Phormidium became highly abundant in 1987 and
afterwards.
There are uncertainties associated with the sampling and quantification of algal
cells. The estimates used in this study were obtained from a single site in the lake,
measured around midday on each occasion. Stochasticities arise from physical
factors such as wind-induced currents, which move algal cells either towards or
away of a sampling site, or vertical positioning of blue-green algal cells in the
water column adjusted by their ability to perform buoyancy. There is also a
likelihood of random and systematic errors in the measurements of physical,
chemical and biological properties of lakes. These factors make predictions of
algal abundance extremely noisy. In addition, algal bloom events are sometimes
only infrequently reflected by the data. That means that the evolutionary algorithm
has to leam an infrequently occurring event from a small noisy dataset.
15.2.1
Data
Initial experiments were conducted using data of 1986 and 1993 as test data, and
data of the remaining eight years as training data. Data of 1986 were typical for a
year with high abundance of the colonial blue-green algae Microcystis spp while
data of 1993 were typical for a year with high abundance of filamentous bluegreen algae such as Oscillatoria spp and Phormidium spp. As Oscillatoria spp and
Phormidium spp are similar in their appearance their data were merged to form a
cumulative dataset of filamentous blue-green algae with highly distinctive bloom
events.
Models induced by machine learning techniques can always be affected by over
training. Over training occurs when the model leams specific characteristics of
patterns, which occur only in the training data but are not truly representative or
general for the domain. To overcome this situation, an explicit knowledge
representation allows the underlying hypothesis of the model to be inspected and
the generalisation of the model to be assessed based on existing knowledge.
Model fitness is an estimate of a model's generalization power, and this fitness
drives the direction of an evolutionary algorithm. When fitness is calculated on the
entire training set it indicates how weIl a model has leamt the examples in the
training set. The ability of a model to generalize can be partly emphasized by only
using a subset of the training set to evaluate fitness. Also either separate validation
sets were used to assess the generalization of evolved models (e.g. Yao and Liu
1997) or information criteria to constrain the complexity of models (e.g. Ghozeil
and FogeI1996).
In order to minimize over training with the current evolutionary algorithm, we
used a random sampie of the training data in each generation to evaluate the
population. Therefore for each generation a training sub set was constructed
consisting of 91 patterns made up by sampling with replacement from the training
Précédent

- 310/410

Suivant