Chapter 14 . Time-Series Prediction of Chlorophyll a in Lakes
269
It is proposed in this study to investigate the effects of hidden layer
configuration, training time, and bagging on the performance of the generic ANN
model structure for predicting algal blooms.
14.2.3
Linear versus Non-Linear Decision Boundaries
One of the primary motivations driving the use of ANNs for modelling
ecosystems is their ability to handle non-linearities in data in a far more graceful
and elegant manner than existing statistical approaches to regression estimation.
This is particularly important in the limnological domain since it is known that it
is characterised by highly complex and non-linear relationships (Whitehead and
Hornberger 1984). However, as discussed above, a trade-off arising from the
ability to model arbitrary decision boundaries is the potential for high model
variance. A pertinent question in light of this trade-off is to ask to what extent do
advantages of lower bias outweigh the ANN achilles heal of high model variance
in a given application.
Frequently, practitioners go to some lengths to justify the use of ANNs by
performing a comparison between a linear regression estimation and the ANN
model estimation. In this study, it is proposed that comparing ANNs with and
without hidden layers is a convenient way to assess the importance of nonlinearity to the model. ANNs without hidden layer neurons (i.e perceptrons) are
constrained to linear decision boundaries, and are functionally and architecturally
similar to a multiple linear regression estimation (Cheng and Titterington 1994).
This study will seek to determine the conditions under which an unconstrained
ANN model (i.e. hidden layer neurons present) succeeds, or fails, to contribute
more to the modelling objectives than an ANN constrained to linear decision
boundaries.
14.3
Implementation of the Generic ANN Aigal Bloom Model
14.3.1
Input Layer Design
An examination of six water quality databases originating from five different lakes
and one river (see table 14.1) revealed four suitable inputs for a generic ANN
model design predicting algal biomass: water temperature, concentrations of
phosphate and nitrate, and water transparency. These inputs were deemed suitable
because i) they are weIl established as important causal
269
It is proposed in this study to investigate the effects of hidden layer
configuration, training time, and bagging on the performance of the generic ANN
model structure for predicting algal blooms.
14.2.3
Linear versus Non-Linear Decision Boundaries
One of the primary motivations driving the use of ANNs for modelling
ecosystems is their ability to handle non-linearities in data in a far more graceful
and elegant manner than existing statistical approaches to regression estimation.
This is particularly important in the limnological domain since it is known that it
is characterised by highly complex and non-linear relationships (Whitehead and
Hornberger 1984). However, as discussed above, a trade-off arising from the
ability to model arbitrary decision boundaries is the potential for high model
variance. A pertinent question in light of this trade-off is to ask to what extent do
advantages of lower bias outweigh the ANN achilles heal of high model variance
in a given application.
Frequently, practitioners go to some lengths to justify the use of ANNs by
performing a comparison between a linear regression estimation and the ANN
model estimation. In this study, it is proposed that comparing ANNs with and
without hidden layers is a convenient way to assess the importance of nonlinearity to the model. ANNs without hidden layer neurons (i.e perceptrons) are
constrained to linear decision boundaries, and are functionally and architecturally
similar to a multiple linear regression estimation (Cheng and Titterington 1994).
This study will seek to determine the conditions under which an unconstrained
ANN model (i.e. hidden layer neurons present) succeeds, or fails, to contribute
more to the modelling objectives than an ANN constrained to linear decision
boundaries.
14.3
Implementation of the Generic ANN Aigal Bloom Model
14.3.1
Input Layer Design
An examination of six water quality databases originating from five different lakes
and one river (see table 14.1) revealed four suitable inputs for a generic ANN
model design predicting algal biomass: water temperature, concentrations of
phosphate and nitrate, and water transparency. These inputs were deemed suitable
because i) they are weIl established as important causal
