160
K. M. Sørensen et al.
step in which a PLS regression model is established using a set of samples measured
twice. One measurement by NIR spectroscopy to collect X and one measurement by
a reference method to collect y. And next a prediction step uses NIR spectroscopy and
the PLS calibration model to predict the value of the response variable for unknown
samples. The benefit of this approach is the ultra-rapid and sustainable measurement
of billions of samples worldwide, but as mentioned in the previous chapter, the
accuracy and reliability of the prediction method will depend on the validation of
the calibration model.
Two examples of PLS calibrations will be shown after discussion on model
validation and error reporting (Section 7.6).
7.5.3 Partial Least Squares Regression—Discriminant
Analysis (PLS-DA)
In NIRS analysis, it is common to have a priori knowledge about the spectral data,
typically from a controlled experimental design (called metadata in the PCA section).
Some metadata are binary or only have discrete levels such as male or female, organic
or conventional, active or placebo, authentic or not, variety 1 or 2, and breed 1 or
2. Such labels can be used actively in regression modeling by introducing a socalled dummy variable that contains the a priori knowledge in the form of a response
variable y vector made of dummy values (typically 0 and 1) distinguishing between
the two different classes. The final assignment of a class to a prediction is done on
a threshold of the predicted dummy y. For instance, if the predicted value is above
0.5, it is assigned to class 1, and if below 0.5, to class 0.
Accordingly, the PLS-DA is a classification method where the dummy variable
is predicted in the best possible way using the information found in the spectral
data [37]. This is closely related to a normal PLS prediction model, where a continuous parameter (e.g., protein level) is predicted from a NIR spectrum, but the main
difference is that PLS-DA solves a classification task. In PLS-DA, the classes are
described in the dummy parameter in the best possible way, providing the best obtainable prediction from a linear combination of the wavelengths, which are weighted
via the regression coefficients in b according to their importance in the prediction
model of the class parameter.
Where a normal PLS model is optimized according to the prediction error (e.g.,
the root mean square error of prediction: RMSEP), the PLS-DA should be optimized based on classification parameters (e.g., rate or percentage of correct and
misclassified samples). PLS-DA is prone to yield overfitted results, and therefore a
thorough validation step (see validation section 7.6) is needed.
K. M. Sørensen et al.
step in which a PLS regression model is established using a set of samples measured
twice. One measurement by NIR spectroscopy to collect X and one measurement by
a reference method to collect y. And next a prediction step uses NIR spectroscopy and
the PLS calibration model to predict the value of the response variable for unknown
samples. The benefit of this approach is the ultra-rapid and sustainable measurement
of billions of samples worldwide, but as mentioned in the previous chapter, the
accuracy and reliability of the prediction method will depend on the validation of
the calibration model.
Two examples of PLS calibrations will be shown after discussion on model
validation and error reporting (Section 7.6).
7.5.3 Partial Least Squares Regression—Discriminant
Analysis (PLS-DA)
In NIRS analysis, it is common to have a priori knowledge about the spectral data,
typically from a controlled experimental design (called metadata in the PCA section).
Some metadata are binary or only have discrete levels such as male or female, organic
or conventional, active or placebo, authentic or not, variety 1 or 2, and breed 1 or
2. Such labels can be used actively in regression modeling by introducing a socalled dummy variable that contains the a priori knowledge in the form of a response
variable y vector made of dummy values (typically 0 and 1) distinguishing between
the two different classes. The final assignment of a class to a prediction is done on
a threshold of the predicted dummy y. For instance, if the predicted value is above
0.5, it is assigned to class 1, and if below 0.5, to class 0.
Accordingly, the PLS-DA is a classification method where the dummy variable
is predicted in the best possible way using the information found in the spectral
data [37]. This is closely related to a normal PLS prediction model, where a continuous parameter (e.g., protein level) is predicted from a NIR spectrum, but the main
difference is that PLS-DA solves a classification task. In PLS-DA, the classes are
described in the dummy parameter in the best possible way, providing the best obtainable prediction from a linear combination of the wavelengths, which are weighted
via the regression coefficients in b according to their importance in the prediction
model of the class parameter.
Where a normal PLS model is optimized according to the prediction error (e.g.,
the root mean square error of prediction: RMSEP), the PLS-DA should be optimized based on classification parameters (e.g., rate or percentage of correct and
misclassified samples). PLS-DA is prone to yield overfitted results, and therefore a
thorough validation step (see validation section 7.6) is needed.
