12 Method Development
279
Fig. 12.1 Generic method development, testing, and validation workflow
in calibration. In addition, there is always error in the determination of the reference
values, and both the accuracy and precision of the reference data should be determined
to understand the impact on the NIR method as the error in the reference method
will impact the performance of the chemometrics model.
Prior to regression, the spectra are often preprocessed to remove unwanted variance and highlight specific information through variable range selection and pretreatment methods. The choice of regression method will depend on the nature of the
signal. The data collected by most spectrometers are highly collinear and will require
the use of regression techniques that can handle the collinearity. The most utilized
methods are based on latent variable extraction techniques [3, 4] (partial least-squares
regression (PLSR), principal component regression (PCR)). These algorithms find
the main directions of variance in the data and correlate them to the reference values
(PCR) or the main directions of covariance between the spectra and the reference
values (PLSR). However, for well-defined systems, Beer’s law can be used through
classical least-squares [4] (CLS) regression or more computationally intensive
techniques such as artificial neural networks (ANN) or support vector machines
(SVM) [5].
After regression, the model stability and performance are internally tested with
cross-validation techniques. Numerous cross-validation approaches have been developed, but, in general, they remove a part of the calibration samples, redevelop the
model without these samples, and subsequently predict the excluded samples. The
operation is repeated until all samples have been used to test the model. Care should be
taken to select the right approach for cross-validation. Leave-one-out cross-validation
is a very simple approach, where only one sample is taken out at a time; but it will tend
to be over-optimistic. Alternatively, if using block cross-validation, an entire source
of variability could be inadvertently removed, and the resulting cross-validation error
would be overly inflated resulting in a loss of confidence that the model will perform
as required. Random block and venetian blinds are alternative sample selection techniques that can provide a more realistic estimate of the model performance. The
selected approach should consider the specificities of the samples in the calibration
set.
279
Fig. 12.1 Generic method development, testing, and validation workflow
in calibration. In addition, there is always error in the determination of the reference
values, and both the accuracy and precision of the reference data should be determined
to understand the impact on the NIR method as the error in the reference method
will impact the performance of the chemometrics model.
Prior to regression, the spectra are often preprocessed to remove unwanted variance and highlight specific information through variable range selection and pretreatment methods. The choice of regression method will depend on the nature of the
signal. The data collected by most spectrometers are highly collinear and will require
the use of regression techniques that can handle the collinearity. The most utilized
methods are based on latent variable extraction techniques [3, 4] (partial least-squares
regression (PLSR), principal component regression (PCR)). These algorithms find
the main directions of variance in the data and correlate them to the reference values
(PCR) or the main directions of covariance between the spectra and the reference
values (PLSR). However, for well-defined systems, Beer’s law can be used through
classical least-squares [4] (CLS) regression or more computationally intensive
techniques such as artificial neural networks (ANN) or support vector machines
(SVM) [5].
After regression, the model stability and performance are internally tested with
cross-validation techniques. Numerous cross-validation approaches have been developed, but, in general, they remove a part of the calibration samples, redevelop the
model without these samples, and subsequently predict the excluded samples. The
operation is repeated until all samples have been used to test the model. Care should be
taken to select the right approach for cross-validation. Leave-one-out cross-validation
is a very simple approach, where only one sample is taken out at a time; but it will tend
to be over-optimistic. Alternatively, if using block cross-validation, an entire source
of variability could be inadvertently removed, and the resulting cross-validation error
would be overly inflated resulting in a loss of confidence that the model will perform
as required. Random block and venetian blinds are alternative sample selection techniques that can provide a more realistic estimate of the model performance. The
selected approach should consider the specificities of the samples in the calibration
set.
