complex, the computational efficiency and the accuracy decreases. In an order to overcome these problems, this study
examines the use of Machine Learning (ML) methods to develop more accurate quantitative methods. In addition, this paper
demonstrates how higher accuracies may be achieved by using an ensemble method that combines the predictions from
multiple methods.
2. EXPERIMENTAL
2.1
Apparatus and materials
The modular instrument used for near-IR Raman spectroscopy has been described previously.
10 All spectra were recorded at
a set interval of 450-1100 cm
-1 (510 data points) and a resolution of ~ 4 cm
-1
. The exposure time was set at 30 seconds for
all samples. Three Raman spectra at different surface locations were recorded for each sample to minimize the effect of
sample morphology. Features in the spectra caused by cosmic rays incident on the detector were manually removed using
the EasyPlot software package (ver. 3.00-7, Spiral Software+MIT). The three spectra were co-added, averaged, and
smoothed over a five-point range. These spectra were then divided by a normalized white light spectrum to correct for
detector response.
Anhydrous D-glucose (BDH), Cocaine hydrochloride and caffeine (Sigma-Aldrich) were reagent grade and were used as
received. The sample set (see Table 1) covered a representative range of concentrations and mixtures with the sample
mixtures (10-30 mg total weight) being made up by mixing known weights of drug and diluent, followed by grinding in an
agate mortar and pestle to ensure sample homogeneity by thorough mixing of components. The mixtures were transferred to
clean stainless steel hexagonal sample holders with an internal diameter of ~ 2 mm and tamped into place.
2.2
Software and Hardware
Chemometric analysis was performed using the Unscrambler (V7.5, CAMO ASA, Trondheim, Norway) multivariate
analysis software package. Neural Network analyses were performed using the Stuttgart Neural Network Simulator
14 and
Genetic Algorithm populations were bred using the GA Playground software
12
. Other software, including the
implementation of the k-NN algorithm (Sec. 3.2), was developed for this project by the authors. Machine Learning analyses
were carried out on a desktop PC and on NUI Galway’s Origin high-performance multi-processor computer.
3. MACHINE LEARNING ANALYSES
3.1
Overview
In this study, the ML analyses have focused on predicting the concentration of cocaine in a sample containing a mixture of
components, by examination of the sample’s Raman spectrum. As mentioned in the Introduction, the analyses involved two
phases: data reduction and prediction. Data reduction involves simplifying the data to improve the accuracy of prediction
and reduce computational effort, as is discussed in detail in Sec. 3.3. Statistical and other transforms may be used for data
reduction, but in this work feature selection was used, which is a simple form of data reduction whereby some of the input
features are selected for use in prediction and the rest are ignored.
Prediction involves building a model of how Raman spectra relate to cocaine concentration, and then using this to predict
the concentration of new samples. Details of the prediction methods are presented in Sec. 3.2. Although prediction and
feature selection are discussed separately below, the two sub-tasks were actually interlinked, as the feature selection was
examines the use of Machine Learning (ML) methods to develop more accurate quantitative methods. In addition, this paper
demonstrates how higher accuracies may be achieved by using an ensemble method that combines the predictions from
multiple methods.
2. EXPERIMENTAL
2.1
Apparatus and materials
The modular instrument used for near-IR Raman spectroscopy has been described previously.
10 All spectra were recorded at
a set interval of 450-1100 cm
-1 (510 data points) and a resolution of ~ 4 cm
-1
. The exposure time was set at 30 seconds for
all samples. Three Raman spectra at different surface locations were recorded for each sample to minimize the effect of
sample morphology. Features in the spectra caused by cosmic rays incident on the detector were manually removed using
the EasyPlot software package (ver. 3.00-7, Spiral Software+MIT). The three spectra were co-added, averaged, and
smoothed over a five-point range. These spectra were then divided by a normalized white light spectrum to correct for
detector response.
Anhydrous D-glucose (BDH), Cocaine hydrochloride and caffeine (Sigma-Aldrich) were reagent grade and were used as
received. The sample set (see Table 1) covered a representative range of concentrations and mixtures with the sample
mixtures (10-30 mg total weight) being made up by mixing known weights of drug and diluent, followed by grinding in an
agate mortar and pestle to ensure sample homogeneity by thorough mixing of components. The mixtures were transferred to
clean stainless steel hexagonal sample holders with an internal diameter of ~ 2 mm and tamped into place.
2.2
Software and Hardware
Chemometric analysis was performed using the Unscrambler (V7.5, CAMO ASA, Trondheim, Norway) multivariate
analysis software package. Neural Network analyses were performed using the Stuttgart Neural Network Simulator
14 and
Genetic Algorithm populations were bred using the GA Playground software
12
. Other software, including the
implementation of the k-NN algorithm (Sec. 3.2), was developed for this project by the authors. Machine Learning analyses
were carried out on a desktop PC and on NUI Galway’s Origin high-performance multi-processor computer.
3. MACHINE LEARNING ANALYSES
3.1
Overview
In this study, the ML analyses have focused on predicting the concentration of cocaine in a sample containing a mixture of
components, by examination of the sample’s Raman spectrum. As mentioned in the Introduction, the analyses involved two
phases: data reduction and prediction. Data reduction involves simplifying the data to improve the accuracy of prediction
and reduce computational effort, as is discussed in detail in Sec. 3.3. Statistical and other transforms may be used for data
reduction, but in this work feature selection was used, which is a simple form of data reduction whereby some of the input
features are selected for use in prediction and the rest are ignored.
Prediction involves building a model of how Raman spectra relate to cocaine concentration, and then using this to predict
the concentration of new samples. Details of the prediction methods are presented in Sec. 3.2. Although prediction and
feature selection are discussed separately below, the two sub-tasks were actually interlinked, as the feature selection was
