Method
RMSEP %
MaxErrP %
RMSEP* %
MaxErrP* %
Partial Least Squares
5.225
17.286
4.421
10.446
Neural Network
5.206
11.628
4.901
10.634
k-Nearest Neighbours
5.837
24.690
4.199
10.203
Ensemble
4.857
17.868
3.891
9.549
Table 3: Performance of individual prediction methods and of an ensemble that averages predictions
of the three methods.
As Table 3 shows, overall performance of the ensemble is better than that of any of the individual members, as shown by
the RMSEP values. Naturally, the MaxErrP of the ensemble cannot be as good as that of the best individual member, since
it averages values from all members. Nonetheless, it is seen that ensemble has a significant ameliorating effect on an
individual bad value such as the MaxErrP of the kNN algorithm.
5. CONCLUSIONS
This paper has investigated the application of Machine Learning techniques for prediction and data reduction to the task of
predicting the concentration of cocaine in solid mixtures, using on Raman spectroscopy. The study has shown that good
results are achievable by Neural Networks and k-Nearest Neighbours, provided that data reduction is used to improve the
dimensionality of the data. In this study, the data reduction took the form of selecting specific wavelengths and discarding
all others. This selection process was optimised by using a Genetic Algorithm, which in combination with the Neural
Network method produced greater prediction accuracy than the Partial Least Squares method. The resulting predictors are
simple, basing their predictions on a small number (less than 20) of data points. Accordingly, after the models have been
built the classifiers operate rapidly, and they could be implemented on hardware for portable probes because they just a
small number of simple mathematical operations (addition and multiplication).
This paper has also demonstrated how an ensemble of different predictors can be used to produce predictions that are better
than any one of the individual predictors, though naturally this comes at the cost of increased computational effort in
constructing multiple predictors.
It appears that all prediction methods considered in this study are fundamentally limited by the lack of information inherent
in a limited test dataset of just 36 samples of various concentrations of cocaine, glucose and caffeine. To achieve
substantially improved prediction accuracies, a much more comprehensive database of samples and their Raman spectra
would be required. In addition, further samples would be necessary for independent verification of the performance of the
prediction methods. Fortunately, in a real-world application, the number of samples available would be continually
expanding as new samples would be analysed on an ongoing basis from law-enforcement seizures and similar sources.
6. ACKNOWLEDGEMENTS
The work was part assisted by the Irish Higher Education Authority, under its Programme for Research in Third Level
Institutions.
RMSEP %
MaxErrP %
RMSEP* %
MaxErrP* %
Partial Least Squares
5.225
17.286
4.421
10.446
Neural Network
5.206
11.628
4.901
10.634
k-Nearest Neighbours
5.837
24.690
4.199
10.203
Ensemble
4.857
17.868
3.891
9.549
Table 3: Performance of individual prediction methods and of an ensemble that averages predictions
of the three methods.
As Table 3 shows, overall performance of the ensemble is better than that of any of the individual members, as shown by
the RMSEP values. Naturally, the MaxErrP of the ensemble cannot be as good as that of the best individual member, since
it averages values from all members. Nonetheless, it is seen that ensemble has a significant ameliorating effect on an
individual bad value such as the MaxErrP of the kNN algorithm.
5. CONCLUSIONS
This paper has investigated the application of Machine Learning techniques for prediction and data reduction to the task of
predicting the concentration of cocaine in solid mixtures, using on Raman spectroscopy. The study has shown that good
results are achievable by Neural Networks and k-Nearest Neighbours, provided that data reduction is used to improve the
dimensionality of the data. In this study, the data reduction took the form of selecting specific wavelengths and discarding
all others. This selection process was optimised by using a Genetic Algorithm, which in combination with the Neural
Network method produced greater prediction accuracy than the Partial Least Squares method. The resulting predictors are
simple, basing their predictions on a small number (less than 20) of data points. Accordingly, after the models have been
built the classifiers operate rapidly, and they could be implemented on hardware for portable probes because they just a
small number of simple mathematical operations (addition and multiplication).
This paper has also demonstrated how an ensemble of different predictors can be used to produce predictions that are better
than any one of the individual predictors, though naturally this comes at the cost of increased computational effort in
constructing multiple predictors.
It appears that all prediction methods considered in this study are fundamentally limited by the lack of information inherent
in a limited test dataset of just 36 samples of various concentrations of cocaine, glucose and caffeine. To achieve
substantially improved prediction accuracies, a much more comprehensive database of samples and their Raman spectra
would be required. In addition, further samples would be necessary for independent verification of the performance of the
prediction methods. Fortunately, in a real-world application, the number of samples available would be continually
expanding as new samples would be analysed on an ongoing basis from law-enforcement seizures and similar sources.
6. ACKNOWLEDGEMENTS
The work was part assisted by the Irish Higher Education Authority, under its Programme for Research in Third Level
Institutions.
