Chapter 11· Input Selection for an Aigal Bloom Model
223
number of generations was exceeded. With a population size of 30 networks, it
was found that optimal ANN models could be determined within 10 generations.
The NGO also has the capability to lock all inputs active and only optimise the
network architecture, wh ich means that it can be used to perform the stepwise
modelling technique. Unless stated otherwise, the default software parameters
were used since the focus is on evaluating the input determination techniques
rather than studying the effect of varying the network's parameters.
11.4.1
Performance Measures and Model Validation
The onset, peak and duration of a bloom or growth event are the three most
important characteristics describing the occurrence of Anabaena spp. The RMSE
is not an ideal measure of fitness but it was considered the most suitable error
measure, as it pI aces greater emphasis on larger forecasting errors. Even if the
RMSEs of several forecasts are similar, the usefulness of the forecasts may differ.
For example, two forecasts may have the same RMSE, but one may forecast
increases that lead the actual event while the other may lag it, making the former
more useful. Therefore, a visual inspection of the plots of actual and predicted
results is important in addition to calculating the RMSE between them. In this
research, the plots of the actual and predicted (4 week forecast) values were
inspected, and the RMSE between them calculated for an arbitrary two-year
validation period spanning November 1992 to November 1994.
11.4.2
Data Division
In this paper, the main objective is to compare different input determination
techniques. To provide a fair comparison between the different models, it is
important that all other modelling factors are held constant and that the models are
tested and validated on data that is statistically representative of the data used in
the training process. This provides the most rigorous test of a model's
performance based on the input selection method since other sources of poor
performance such as attempting to validate the model on data outside the range
used in training are effectively eliminated. In addition, the ASCE Task
Committee on Application of Artificial Neural Networks in Hydrology (2000)
observed that "an optimal data set for training would be one that fully represents
the modeling domain and has the minimum number of data pairs in training". To
achieve this, the data were divided into training and testing sets using a SOM.
The validation data were combined with the remaining data and clustered using
the SOM. Once the clusters were formed, two data records from each cluster
containing validation data were sampled (i.e. one for each of the training and
testing sets). In the instance that a cluster only contained one record other than the
validation record, then this record was placed in the training set. The advantage of
using the SOM data division technique is that it employs the minimum number of
223
number of generations was exceeded. With a population size of 30 networks, it
was found that optimal ANN models could be determined within 10 generations.
The NGO also has the capability to lock all inputs active and only optimise the
network architecture, wh ich means that it can be used to perform the stepwise
modelling technique. Unless stated otherwise, the default software parameters
were used since the focus is on evaluating the input determination techniques
rather than studying the effect of varying the network's parameters.
11.4.1
Performance Measures and Model Validation
The onset, peak and duration of a bloom or growth event are the three most
important characteristics describing the occurrence of Anabaena spp. The RMSE
is not an ideal measure of fitness but it was considered the most suitable error
measure, as it pI aces greater emphasis on larger forecasting errors. Even if the
RMSEs of several forecasts are similar, the usefulness of the forecasts may differ.
For example, two forecasts may have the same RMSE, but one may forecast
increases that lead the actual event while the other may lag it, making the former
more useful. Therefore, a visual inspection of the plots of actual and predicted
results is important in addition to calculating the RMSE between them. In this
research, the plots of the actual and predicted (4 week forecast) values were
inspected, and the RMSE between them calculated for an arbitrary two-year
validation period spanning November 1992 to November 1994.
11.4.2
Data Division
In this paper, the main objective is to compare different input determination
techniques. To provide a fair comparison between the different models, it is
important that all other modelling factors are held constant and that the models are
tested and validated on data that is statistically representative of the data used in
the training process. This provides the most rigorous test of a model's
performance based on the input selection method since other sources of poor
performance such as attempting to validate the model on data outside the range
used in training are effectively eliminated. In addition, the ASCE Task
Committee on Application of Artificial Neural Networks in Hydrology (2000)
observed that "an optimal data set for training would be one that fully represents
the modeling domain and has the minimum number of data pairs in training". To
achieve this, the data were divided into training and testing sets using a SOM.
The validation data were combined with the remaining data and clustered using
the SOM. Once the clusters were formed, two data records from each cluster
containing validation data were sampled (i.e. one for each of the training and
testing sets). In the instance that a cluster only contained one record other than the
validation record, then this record was placed in the training set. The advantage of
using the SOM data division technique is that it employs the minimum number of
