0
0.05
0.1
0.15
0.2
0.25
0.3
0.35
0
100
200
300
400
500
Data Point No.
Intensity
100% Cocaine
Maxima
GA-kNN
GA-NNet
Figure 1: Results of attribute selection procedures.
4.2
Results of Analyses
Table 2 summarises the performance of the kNN and neural network methods, when combined with the different attribute
selection strategies, and also includes the performance of the Partial Least Squares technique as applied to this dataset
previously by the authors
11
. For each combination of prediction method and attribute selection technique, the table lists the
number of attributes selected, the root mean squared error of prediction (RMSEP) and the absolute maximum error in
prediction (MaxErrP). In all cases, the prediction procedure was standard leave-one-out cross-validation: for each sample in
turn, that sample was removed from the set and the remainder was used to build a model, which was then used to predict the
concentration of cocaine in the sample that had been removed. For the NN and PLS methods, calibration statistics (RMSEC
and MaxErrC) are also listed: these are found from using all samples together to build a model, and then using the model to
predict the concentration of each sample in turn. In the NN case, this is essentially the training error of the method. For the
kNN algorithm, calibration statistics are not meaningful – for the 1-neighbour case, if the sample to be predicted is
available, the kNN algorithm will have zero error. As is typical of neural networks, it can be seen in rows 5-7 that
calibration statistics are much better than prediction statistics, because of the way the network can represent nonlinear data
accurately.
Examining Table 2, it is seen that the kNN and NN methods both benefit from careful selection of attributes, as reflected in
reduced values for RMSEP and MaxErrP when GA-based selection is used compared to when no selection or simple
maxima-based selection is used. In the best case, the NN outperforms the PLS method both in terms of root mean squared
error and maximum error. However, the performance of the kNN method is not quite as good. Overall, the best-case
performance of all methods are quite similar; it appears that all methods are fundamentally limited by a lack of information
Précédent

- 7/11

Suivant