274
H. Wilson . F. Recknagel
14.3.4
Model Assessment
The standard practice of assessing how weIl ANN models generalise to population
data is to analyse predictions made on an independent subsampIe of data not
included in the training set. This validation approach has been utilised for all
previous ANN applications to modelling limnological variables.
In this
application it was decided to extend this approach in several ways.
One of the problems of the so-called "split plot" approach is that it causes a
dilemma with respect to the efficient use of data in the context of data constrained
applications. Achieving a test set representation capable of producing a reliable
estimation of model performance on new data means compromising training set
representation. Similarly, maximising training set representation reduces the
testing opportunities.
A simple solution to this dilemma involves resampling, where multiple models
are trained using resampled training and test sets. The performance of the model
is then estimated on the combined test set results. Common resample techniques
include leave-l-out (i.e each record in a sampIe sized n used for testing in turn),
and k-fold-crossvalidation (k groups of records in the sampie are left out for
testing in turn) (Weiss and Kulikowski 1991). Wilson and Recknagel (1997)
demonstrated the use of a 10-fold-crossvalidation technique for assessment of
ANN models predicting growth dynamics of 8 species of phytoplankton in Lake
Kasumigaura, Japan.
Another resampling method of model evaluation involves the use of bootstrap
sampies for training, where for a sampIe size of n, n records are independently
sampled with replacement. The remaining unsampled records are then used for
model assessment. The principle advantage of this approach is that it becomes
possible to determine the effect of sampling variability on the model. However,
Weiss and Kulikowski (1991) point out that on small sampIes, this approach is
pessimistically biased compared to leave-one-out crossvalidation because on
average, only 63.2% of the total sampIe is used for training.
In this application, it was elected to use the bootstrap resampling method to
select training and test set data. The process of bootstrap sampling, training and
testing was carried out for 100 times for each of the model structures in a factorial
design. Model performance was then assessed in 3 ways:
1. The test set performance of each of the bootstrap models was evaluated.
2. The minimum, maximum, median, mean and interquartile ranges of the test set
predictions for each value in the sampIe was calculated. RMS error values
were calculated for the mean predictions.
3. Visual assessment of the mean model predictions versus the observed algal
abundance measures was performed.
The second and third model assessment procedures evaluate the bootstrap
aggregate, or bagged model.
H. Wilson . F. Recknagel
14.3.4
Model Assessment
The standard practice of assessing how weIl ANN models generalise to population
data is to analyse predictions made on an independent subsampIe of data not
included in the training set. This validation approach has been utilised for all
previous ANN applications to modelling limnological variables.
In this
application it was decided to extend this approach in several ways.
One of the problems of the so-called "split plot" approach is that it causes a
dilemma with respect to the efficient use of data in the context of data constrained
applications. Achieving a test set representation capable of producing a reliable
estimation of model performance on new data means compromising training set
representation. Similarly, maximising training set representation reduces the
testing opportunities.
A simple solution to this dilemma involves resampling, where multiple models
are trained using resampled training and test sets. The performance of the model
is then estimated on the combined test set results. Common resample techniques
include leave-l-out (i.e each record in a sampIe sized n used for testing in turn),
and k-fold-crossvalidation (k groups of records in the sampie are left out for
testing in turn) (Weiss and Kulikowski 1991). Wilson and Recknagel (1997)
demonstrated the use of a 10-fold-crossvalidation technique for assessment of
ANN models predicting growth dynamics of 8 species of phytoplankton in Lake
Kasumigaura, Japan.
Another resampling method of model evaluation involves the use of bootstrap
sampies for training, where for a sampIe size of n, n records are independently
sampled with replacement. The remaining unsampled records are then used for
model assessment. The principle advantage of this approach is that it becomes
possible to determine the effect of sampling variability on the model. However,
Weiss and Kulikowski (1991) point out that on small sampIes, this approach is
pessimistically biased compared to leave-one-out crossvalidation because on
average, only 63.2% of the total sampIe is used for training.
In this application, it was elected to use the bootstrap resampling method to
select training and test set data. The process of bootstrap sampling, training and
testing was carried out for 100 times for each of the model structures in a factorial
design. Model performance was then assessed in 3 ways:
1. The test set performance of each of the bootstrap models was evaluated.
2. The minimum, maximum, median, mean and interquartile ranges of the test set
predictions for each value in the sampIe was calculated. RMS error values
were calculated for the mean predictions.
3. Visual assessment of the mean model predictions versus the observed algal
abundance measures was performed.
The second and third model assessment procedures evaluate the bootstrap
aggregate, or bagged model.
