the corresponding digit in the bit-string was 1 or 0. The fitness of each individual was assessed by measuring the
performance of kNN with the indicated set of attributes selected, and giving a penalty based on the number of attribute
selected, to encourage the development of individuals that selected the best set of attributes and to favour the selection of as
few attributes as possible. Populations were bred using the GA Playground software
12
, a general-purpose Genetic Algorithm
toolkit implemented in Java, which the authors interfaced to their own kNN software. Repeated runs were carried out with
various population sizes, crossover rates, mutation rates, and penalties.
In the second set of GA experiments, a neural network was the target learning algorithm. Because of the relative slowness
of neural network training, the fitness function was based on the sum of the squared error in training, since full crossvalidation for the 36 samples would have taken 36 times longer. Likewise, the 510 attributes per sample were reduced to a
representative set of less than 10%, comprising local maxima, local minima and some intermediate points. In addition,
because comparisons across networks with different sets of input attributes were being performed, a fixed number of hidden
nodes was used for all networks and all networks were trained for the name number of epochs. Having a penalty term to
encourage solutions with fewer attributes was not found to be beneficial in the neural network case.
3.4
Ensembles
Ensemble methods are currently an active area of research within machine learning. An ensemble is a learning algorithm
that uses a set of predictors and combines their predictions through a voting scheme to arrive at a decision. Ensembles
produce more accurate results than their individual members provided that they are accurate (i.e. their performance is better
than random) and diverse (i.e. different members make different errors on new data)
17
. A good overview is given by
Dietterich
18
. Ensembles are usually constructed by producing variations on a single classifier, for example by training
several neural networks using different feature subsets for each. In this study, however, completely different prediction
methods are used for constructing the members of the ensemble.
4. RESULTS & DISCUSSION
4.1
Results of Feature Selection
Figure 2 illustrates the effect of the feature selection features on the Raman spectrum of a pure cocaine sample. Without any
feature selection, the full spectrum of 510 points is used. If the maxima are used, this reduces the data set to 17 points per
sample, corresponding to the local maxima of the pure cocaine sample (even though these may not be maxima in other
samples). The plot also shows the results of the GA optimisation when the kNN algorithm and the NN algorithm are each
used as targets for the optimisation procedure. Lines connect the points to make them easier to identify. These two results
provide an interesting contrast with each other: for kNN, the optimal solution is a set of just four points, three of which are
clustered around the largest peak at 996 cm
-1
. For the neural network, on the other hand, the optimal solution does not use
points on that peak at all, instead being based on other local maxima and local minima of the spectrum.
performance of kNN with the indicated set of attributes selected, and giving a penalty based on the number of attribute
selected, to encourage the development of individuals that selected the best set of attributes and to favour the selection of as
few attributes as possible. Populations were bred using the GA Playground software
12
, a general-purpose Genetic Algorithm
toolkit implemented in Java, which the authors interfaced to their own kNN software. Repeated runs were carried out with
various population sizes, crossover rates, mutation rates, and penalties.
In the second set of GA experiments, a neural network was the target learning algorithm. Because of the relative slowness
of neural network training, the fitness function was based on the sum of the squared error in training, since full crossvalidation for the 36 samples would have taken 36 times longer. Likewise, the 510 attributes per sample were reduced to a
representative set of less than 10%, comprising local maxima, local minima and some intermediate points. In addition,
because comparisons across networks with different sets of input attributes were being performed, a fixed number of hidden
nodes was used for all networks and all networks were trained for the name number of epochs. Having a penalty term to
encourage solutions with fewer attributes was not found to be beneficial in the neural network case.
3.4
Ensembles
Ensemble methods are currently an active area of research within machine learning. An ensemble is a learning algorithm
that uses a set of predictors and combines their predictions through a voting scheme to arrive at a decision. Ensembles
produce more accurate results than their individual members provided that they are accurate (i.e. their performance is better
than random) and diverse (i.e. different members make different errors on new data)
17
. A good overview is given by
Dietterich
18
. Ensembles are usually constructed by producing variations on a single classifier, for example by training
several neural networks using different feature subsets for each. In this study, however, completely different prediction
methods are used for constructing the members of the ensemble.
4. RESULTS & DISCUSSION
4.1
Results of Feature Selection
Figure 2 illustrates the effect of the feature selection features on the Raman spectrum of a pure cocaine sample. Without any
feature selection, the full spectrum of 510 points is used. If the maxima are used, this reduces the data set to 17 points per
sample, corresponding to the local maxima of the pure cocaine sample (even though these may not be maxima in other
samples). The plot also shows the results of the GA optimisation when the kNN algorithm and the NN algorithm are each
used as targets for the optimisation procedure. Lines connect the points to make them easier to identify. These two results
provide an interesting contrast with each other: for kNN, the optimal solution is a set of just four points, three of which are
clustered around the largest peak at 996 cm
-1
. For the neural network, on the other hand, the optimal solution does not use
points on that peak at all, instead being based on other local maxima and local minima of the spectrum.
