7 NIR Data Exploration and Regression by Chemometrics—A Primer
153
The second loading (cyan) is the structure that describes the second largest amount
of variation in the dataset where this second loading vector also has to have the
constrained of being orthogonal to the first loading. Again, the difference in the
samples is evident only from the score value, which is −0.01 for sample #43.
The part of the variation in the dataset that is not described by the first two principal
components is shown in the residuals in the right most column of Fig. 7.16. The
residuals are specific for each sample and may for example be used to detect aberrant
patterns in single measurements. Note the y-axis of the residuals: The numerical
values fluctuate within ±0.002. These values can be directly comparable to the
variation in the mean-centered spectral data (Fig. 7.16), which varies between 0.1
and 0.7.
By comparing the size of the residuals with the variation of the mean-centered
data, the explained variance can be calculated for each principal component. In this
case, the first component (PC1) explains 88.0% of the initial total variation, the
second component (PC2) explains 11.6% of the remaining variation, and overall the
two components thus explain 99.6% of the variation in the dataset.
Plotting all 231 score values for the first principal component against the corresponding values for the second component yields a score scatter plot (Fig. 7.17) in
which each point represents a NIR spectrum with originally 700 variables. In the
given case, sample #43 can be seen in the coordinate system with the coordinates
(−0.79; 0.01), sample #107 at (0.17; 0.11), sample #224 at (0.52; -0.25) and so on
for the remaining samples.
As shown by the example, PCA is a good tool for exploratory data analysis of
highly colinear data as often seen in spectroscopy. As a result, one can observe the
behavior and characteristics of single samples and study, which wavelength ranges
are important for the similarity or difference between samples. PCA can be perceived
as a “reverse” Lambert–Beer model: The model estimates latent spectra (loadings)
and determines the (pseudo)concentrations of these in the samples (scores) from the
measured spectra. For spectroscopists, the disadvantage is the tricky interpretation
of the loadings, which are not pure analyte spectra representative of the underlying
chemistry. The main take-home message of this PCA application is that samples
which are close to each other in composition are also close in the score plot; i.e.,
biological replicates in your dataset, e.g., should be found close to each other. For
Dataset 2, we observe the experimental design and it is characteristic for PCA that the
score plot of this 3-component mixture is completely described by two components
(the chemical rank is 2 because of closure where the three concentrations add up to
100%). This is in contrast to the MCR model (Fig. 7.14) which models one component
for each chemical component.
7.4.4 PCA for Outlier Detection
The ability of the PCA to reveal the behavior and characteristics of single samples as
part of the complete sample set makes it a powerful tool in the detection of outliers.
Précédent

- 158/586

Suivant