220
G.J. Bowden . G.C. Dandy . H.R. Maier
clusters were formed from the training data. Onee the clusters are formed, one
input from each cluster is sampled and used in the final subset of input variables.
Principal Component Analysis (PCA)
When many potential variables are available, PCA can be used to reduce the
dimensionality of the input data set. By using principal components (PCs), the
variables can be transformed to a new, smaller set of variables, which capture
most of the information in the original data set. This is achieved by computing
factors (new variables) as linear combinations of the old variables. The weights
are selected in such a way as to ensure that some optimality criterion is maximised
(Masters 1995).
To commence the procedure, a single linear combination of all variables is
sought such that the majority of the variation in the training set is captured. After
a single dominant factor is found, it is then necessary to find a second factor that
captures the remaining information not explained by the first factor. The second
PC is chosen such that it is orthogonal to the first. Next a third factor is sought
that best captures the remaining information and is orthogonal to the first and
second. The process continues until all the variance in the data set is accounted
for. If there are p input variables of interest, it is hoped that m, where m< different PCs will account for most of the variation in x. As the interrelations
among the variables increases, the proportion of variance explained by the first
few components increases. Hence, it is common for most of the important
information to be concentrated in the first few principal components, with the
system noise falling mostly in later components that can be discarded (Masters
1995). However, it is possible, but usually unlikely, that by discarding the later
components, some important information may be lost (Jolliffe 1986; Masters
1995). An added advantage of PCA is that eaeh of the computed factors are
independent of each other and there is no redundancy in the information that they
contain (Masters 1995).
If the input variables have different units, it may be necessary to normalise the
data. PCA attempts to capture variation. Numerically, a variation in flow
between 10,000 and 20,000 ML/day is much greater than a variation in river level
between 1.0 and 5.0 m. However, the effect of each of these variables on the
system being investigated may be rather similar and the information content of the
flow data is not inherently greater. Hence, by normalising the data, all variables
are put on an equal basis in the analysis.
11.2.2
Supervised Input Determination
It is important to note that substantial variation amongst the input variables does
not necessarily imply any relationship with the output variable. Hence, after the
input dimensionality has been reduced using the unsupervised procedures (a priori
identification, SOM and PCA), supervised procedures must ultimately decide
upon which variables have the most significant impact on the ANN's forecasting
ability.
G.J. Bowden . G.C. Dandy . H.R. Maier
clusters were formed from the training data. Onee the clusters are formed, one
input from each cluster is sampled and used in the final subset of input variables.
Principal Component Analysis (PCA)
When many potential variables are available, PCA can be used to reduce the
dimensionality of the input data set. By using principal components (PCs), the
variables can be transformed to a new, smaller set of variables, which capture
most of the information in the original data set. This is achieved by computing
factors (new variables) as linear combinations of the old variables. The weights
are selected in such a way as to ensure that some optimality criterion is maximised
(Masters 1995).
To commence the procedure, a single linear combination of all variables is
sought such that the majority of the variation in the training set is captured. After
a single dominant factor is found, it is then necessary to find a second factor that
captures the remaining information not explained by the first factor. The second
PC is chosen such that it is orthogonal to the first. Next a third factor is sought
that best captures the remaining information and is orthogonal to the first and
second. The process continues until all the variance in the data set is accounted
for. If there are p input variables of interest, it is hoped that m, where m< different PCs will account for most of the variation in x. As the interrelations
among the variables increases, the proportion of variance explained by the first
few components increases. Hence, it is common for most of the important
information to be concentrated in the first few principal components, with the
system noise falling mostly in later components that can be discarded (Masters
1995). However, it is possible, but usually unlikely, that by discarding the later
components, some important information may be lost (Jolliffe 1986; Masters
1995). An added advantage of PCA is that eaeh of the computed factors are
independent of each other and there is no redundancy in the information that they
contain (Masters 1995).
If the input variables have different units, it may be necessary to normalise the
data. PCA attempts to capture variation. Numerically, a variation in flow
between 10,000 and 20,000 ML/day is much greater than a variation in river level
between 1.0 and 5.0 m. However, the effect of each of these variables on the
system being investigated may be rather similar and the information content of the
flow data is not inherently greater. Hence, by normalising the data, all variables
are put on an equal basis in the analysis.
11.2.2
Supervised Input Determination
It is important to note that substantial variation amongst the input variables does
not necessarily imply any relationship with the output variable. Hence, after the
input dimensionality has been reduced using the unsupervised procedures (a priori
identification, SOM and PCA), supervised procedures must ultimately decide
upon which variables have the most significant impact on the ANN's forecasting
ability.
