Chapter 11 . Input Selection for an Aigal Bloorn Model
217
In this paper, as a first step in preprocessing the input variables, unsupervised
techniques (Section 11.2.1) have been used to reduce the dimensionality of the
input space. These techniques are the Self-Organizing Map (SOM) and principal
component analysis (PCA). They are compared with the commonly employed
approach of using apriori knowledge of the system to be modelIed. Once
redundant inputs have been removed, the input subsets obtained from each
unsupervised method are further refined using two supervised input selection
methods (Section 11.2.2). These are a hybrid genetic algorithm (GA) and ANN
(GA-ANN) and a stepwise ANN modelling procedure. The input determination
methods have been used for selecting the optimal subset of input variables for
forecasting the concentration of Anabaena spp. in the River Murray at Morgan, 4
weeks in advance. The model inputs obtained using each method are compared
and each of the six sets of inputs have been used to develop ANN models for
forecasting Anabaena spp. in the River Murray at Morgan. The ANNs'
performance on an independent validation set was used to assess the adequacy of
each method of input determination.
11.2
Methods
11.2.1
Unsupervised Input Pre-Processing
Apriori identification
In a typical ANN forecasting application, the modeller collects all time series data,
subject to availability, that is likely to have an influence on the output variable.
Obviously, some knowledge of the system is assumed in determining this set of
candidate input variables. However, the data set, although comprehensive, is
likely to contain some redundant information. An unsupervised approach to
reduce the dimensionality of the input data is to use expert knowledge of the
system being modelIed. In this way the set of all variables likely to influence the
output variable can be reduced to a subset of only those variables most likely to
have a significant influence. Expert knowledge can also be used to select the
maximum lag of each variable chosen. To aid in this task, it is possible to make
use of time series plots of each potential input variable and the output variable.
Inspecting the data plots gives a visual indication of any potential relationship that
may ex ist between the input and output variable.
Apriori identification is widely used in many ANN applications and since it is
dependent on an expert's knowledge, it is very subjective and case dependent.
That is why the two analytical procedures (the SOM and PCA) are also being
considered.
217
In this paper, as a first step in preprocessing the input variables, unsupervised
techniques (Section 11.2.1) have been used to reduce the dimensionality of the
input space. These techniques are the Self-Organizing Map (SOM) and principal
component analysis (PCA). They are compared with the commonly employed
approach of using apriori knowledge of the system to be modelIed. Once
redundant inputs have been removed, the input subsets obtained from each
unsupervised method are further refined using two supervised input selection
methods (Section 11.2.2). These are a hybrid genetic algorithm (GA) and ANN
(GA-ANN) and a stepwise ANN modelling procedure. The input determination
methods have been used for selecting the optimal subset of input variables for
forecasting the concentration of Anabaena spp. in the River Murray at Morgan, 4
weeks in advance. The model inputs obtained using each method are compared
and each of the six sets of inputs have been used to develop ANN models for
forecasting Anabaena spp. in the River Murray at Morgan. The ANNs'
performance on an independent validation set was used to assess the adequacy of
each method of input determination.
11.2
Methods
11.2.1
Unsupervised Input Pre-Processing
Apriori identification
In a typical ANN forecasting application, the modeller collects all time series data,
subject to availability, that is likely to have an influence on the output variable.
Obviously, some knowledge of the system is assumed in determining this set of
candidate input variables. However, the data set, although comprehensive, is
likely to contain some redundant information. An unsupervised approach to
reduce the dimensionality of the input data is to use expert knowledge of the
system being modelIed. In this way the set of all variables likely to influence the
output variable can be reduced to a subset of only those variables most likely to
have a significant influence. Expert knowledge can also be used to select the
maximum lag of each variable chosen. To aid in this task, it is possible to make
use of time series plots of each potential input variable and the output variable.
Inspecting the data plots gives a visual indication of any potential relationship that
may ex ist between the input and output variable.
Apriori identification is widely used in many ANN applications and since it is
dependent on an expert's knowledge, it is very subjective and case dependent.
That is why the two analytical procedures (the SOM and PCA) are also being
considered.
