226
GJ. Bowden . G.C. Dandy . H.R. Maier
Graphical comparisons of the 4-week forecasts of the concentration of Anabaena
spp. in the River Murray at Morgan for the validation period obtained using model
1 and model 2 are shown in Figures 11.2 and 11.3, respectively. Both of these
models were developed using apriori knowledge as the unsupervised input
processing technique. Model 1 used a hybrid GA-ANN for the supervised input
selection, whereas model 2 used the stepwise ANN modelling procedure. The
precision of the cell count data can range from ±20% to ±70% depending on the
number of Anabaena trichomes counted (M.D. Burch, personal communication).
Hence, error bars were conservatively set at ±30% and included in Figures 11.2,
11.3 and 11.4. From Figure 11.2, it can be seen that modell is able to predict the
onset and duration of the two large growth events but underestimates the
magnitude of the second large peak. In Figure 11.3, it can be seen that model 2
was unsuccessful at predicting the onset of the two large growth events and
instead provided a forecast which led the actual events. However, model 2 was
successful at predicting the duration and relative magnitude of the two growth
events. Model 2 only used flow, lags 1-10 weeks and turbidity, lags 1-10 weeks
as input variables, whereas Model 1 made use of 7 different variables including
past lags of flow, turbidity, Anabaena spp., nutrient data, colour, temperature and
pH (Table 11.2). Therefore, other variables are required in addition to flow and
turbidity, in order to forecast the onset of the Anabaena spp. growth events. It is
also important to note that the significant lags selected by the GA-ANN were not
intuitive. The superior performance shown by modelion the validation set may
have resulted from the GA-ANN finding a synergistic combination of input
variables and lags. Such a combination of inputs would be impossible to
determine using the stepwise ANN modelling procedure. The stepwise ANN
procedure is more likely to overlook combinations of interdependent variables,
which may together carry significant information.
A plot of the 4-week forecasts for the validation period obtained using model 3 is
shown in Figure 11.4. Model 3 used PCA as the unsupervised processing
technique and a GA-ANN for the supervised input selection. From Figure 11.4, it
can be seen that model 3 picked up the general shape of the two large peaks of
Anabaena spp. but underestimated the duration and relative magnitude of the
second peak. In general, model 3 was unable to perform as weH on the validation
data as model 1 (the model developed using apriori knowledge and the GAANN) (Figure 11.2).
The RMSEs of the 4-week forecasts for each of the models developed are
shown in Table 11.3. Looking at the performance on the validation set, it can be
seen that the models developed using apriori knowledge performed significantly
better than the models developed using PCA and the SOM as unsupervised input
processing techniques. This suggests that, where available, expert knowledge on
the system being modelled provides a suitable means for removing redundant
input variables and lags of variables. However, this technique is very subjective
and dependent on the case study under investigation. In addition, this type of
expert knowledge is often unavailable and the only alternative is to proceed with
an analytical technique. Based on the performance on the validation set, PCA and
the SOM technique both provided an equaHy suitable means of reducing the
GJ. Bowden . G.C. Dandy . H.R. Maier
Graphical comparisons of the 4-week forecasts of the concentration of Anabaena
spp. in the River Murray at Morgan for the validation period obtained using model
1 and model 2 are shown in Figures 11.2 and 11.3, respectively. Both of these
models were developed using apriori knowledge as the unsupervised input
processing technique. Model 1 used a hybrid GA-ANN for the supervised input
selection, whereas model 2 used the stepwise ANN modelling procedure. The
precision of the cell count data can range from ±20% to ±70% depending on the
number of Anabaena trichomes counted (M.D. Burch, personal communication).
Hence, error bars were conservatively set at ±30% and included in Figures 11.2,
11.3 and 11.4. From Figure 11.2, it can be seen that modell is able to predict the
onset and duration of the two large growth events but underestimates the
magnitude of the second large peak. In Figure 11.3, it can be seen that model 2
was unsuccessful at predicting the onset of the two large growth events and
instead provided a forecast which led the actual events. However, model 2 was
successful at predicting the duration and relative magnitude of the two growth
events. Model 2 only used flow, lags 1-10 weeks and turbidity, lags 1-10 weeks
as input variables, whereas Model 1 made use of 7 different variables including
past lags of flow, turbidity, Anabaena spp., nutrient data, colour, temperature and
pH (Table 11.2). Therefore, other variables are required in addition to flow and
turbidity, in order to forecast the onset of the Anabaena spp. growth events. It is
also important to note that the significant lags selected by the GA-ANN were not
intuitive. The superior performance shown by modelion the validation set may
have resulted from the GA-ANN finding a synergistic combination of input
variables and lags. Such a combination of inputs would be impossible to
determine using the stepwise ANN modelling procedure. The stepwise ANN
procedure is more likely to overlook combinations of interdependent variables,
which may together carry significant information.
A plot of the 4-week forecasts for the validation period obtained using model 3 is
shown in Figure 11.4. Model 3 used PCA as the unsupervised processing
technique and a GA-ANN for the supervised input selection. From Figure 11.4, it
can be seen that model 3 picked up the general shape of the two large peaks of
Anabaena spp. but underestimated the duration and relative magnitude of the
second peak. In general, model 3 was unable to perform as weH on the validation
data as model 1 (the model developed using apriori knowledge and the GAANN) (Figure 11.2).
The RMSEs of the 4-week forecasts for each of the models developed are
shown in Table 11.3. Looking at the performance on the validation set, it can be
seen that the models developed using apriori knowledge performed significantly
better than the models developed using PCA and the SOM as unsupervised input
processing techniques. This suggests that, where available, expert knowledge on
the system being modelled provides a suitable means for removing redundant
input variables and lags of variables. However, this technique is very subjective
and dependent on the case study under investigation. In addition, this type of
expert knowledge is often unavailable and the only alternative is to proceed with
an analytical technique. Based on the performance on the validation set, PCA and
the SOM technique both provided an equaHy suitable means of reducing the
