7 NIR Data Exploration and Regression by Chemometrics—A Primer
157
7.4.6 Outro
PCA is implemented in all chemometric software packages, and it cannot be emphasized enough that performing a PCA on a recorded spectral dataset is a prerequisite
to understand the variability in the data. This also includes the effect of different
spectral pre-processing methods. Usually, the more focused and systematic the score
plot is, the better and higher the amount of explained variance from the first PCs
(so-called parsimony), the better. It should be noted here that PCA score plots do not
change by validation, but the explained variance of the principal components does.
7.5 Calibration by Partial Least Squares (PLS) Regression
PLS is one of the strongest regression methods invented.
It works where Multiple Linear Regression fails!
The task of multivariate calibration is to find a predictive model that relates the NIR
instrumental response space to the analyte concentration space. Here, the purpose of
calibration is to model analyte concentrations, y, as linear combinations of absorption spectra X. Next, analyte concentrations in future samples can be predicted based
on the absorption spectra only. Where PCA represents an untargeted and unsupervised data exploration, partial least squares (PLS) regression [32] is the targeted and
supervised method par excellence.
7.5.1 Regression with Principal Components
When data matrix X is decomposed with a PCA, it is represented in a model space,
represented by the principal components describing the systematic variance. If a
relationship can be found between this model space of the data and an independent
or reference variable, the independent variable can be explained in terms of the
observed dependent variance and hence a regression can be made. This process is
called calibration.
A classical regression extension of PCA is known as principal component regression (PCR) [33]. Having projected a data matrix X into a PCA model defined by
loadings P, resulting in scores T, a regression toward a dependent y can be made via
the regression vector b in the model space:
X = T · P
T
+ E
y = T · b
T
+ q
(7.25)
157
7.4.6 Outro
PCA is implemented in all chemometric software packages, and it cannot be emphasized enough that performing a PCA on a recorded spectral dataset is a prerequisite
to understand the variability in the data. This also includes the effect of different
spectral pre-processing methods. Usually, the more focused and systematic the score
plot is, the better and higher the amount of explained variance from the first PCs
(so-called parsimony), the better. It should be noted here that PCA score plots do not
change by validation, but the explained variance of the principal components does.
7.5 Calibration by Partial Least Squares (PLS) Regression
PLS is one of the strongest regression methods invented.
It works where Multiple Linear Regression fails!
The task of multivariate calibration is to find a predictive model that relates the NIR
instrumental response space to the analyte concentration space. Here, the purpose of
calibration is to model analyte concentrations, y, as linear combinations of absorption spectra X. Next, analyte concentrations in future samples can be predicted based
on the absorption spectra only. Where PCA represents an untargeted and unsupervised data exploration, partial least squares (PLS) regression [32] is the targeted and
supervised method par excellence.
7.5.1 Regression with Principal Components
When data matrix X is decomposed with a PCA, it is represented in a model space,
represented by the principal components describing the systematic variance. If a
relationship can be found between this model space of the data and an independent
or reference variable, the independent variable can be explained in terms of the
observed dependent variance and hence a regression can be made. This process is
called calibration.
A classical regression extension of PCA is known as principal component regression (PCR) [33]. Having projected a data matrix X into a PCA model defined by
loadings P, resulting in scores T, a regression toward a dependent y can be made via
the regression vector b in the model space:
X = T · P
T
+ E
y = T · b
T
+ q
(7.25)
