154
K. M. Sørensen et al.
Outliers can be defined as samples that have a different variable pattern compared
to other samples in the dataset. When many variables have been measured (spectral
data), it can often be difficult to find such patterns via direct inspection, i.e., plotting,
of measurement data. Using PCA, it is possible to find deviating spectra (outliers)
using relevant graphical images that lead directly back to measurement data. In PCA,
the pattern or relationship between all variables is analyzed and handled through the
calculated loadings (P). The basic assumption in PCA is that all samples can be
described with the same set of loadings. Samples for which this does not apply will
have a variable pattern that differs from all other samples in the dataset.
In PCA, there are two distance measures that differentiate how much a sample
differs from the rest of the dataset: the size of the residual and Hotelling’s T 2 . The
residual of a sample can be calculated directly from the residual E matrix. For a given
sample, the square sum of all the elements of the corresponding row in E is calculated
(see, e.g., Fig. 7.16). A sample with higher residual variance will have a pattern or
variation in the original data that is not similar to the remaining samples. The second
most important distance measure is based on the score values (T). The distance to
the center of a sample in the score space can be calculated using Hotelling’s T 2 ,
which considers the covariance in the data. Combining the two distance measures in
a scatter plot, i.e., the residual variance and Hotelling’s T 2 provide the most important
diagnostic plots in PCA.
It is of fundamental scientific importance to be able to efficiently identify outliers
as they may represent new discoveries with completely new functionalities or, as
a contrast, identification of samples that ruins the models. In chemometric modeling,
outliers are undesirable because they are included in the estimation of model parameters. Thus, the PCA model must be recalculated, when one or more samples are
characterized as outliers and discarded. It is thus an iterative process to characterize
and eliminate outliers. This is easily done in modern chemometric software where
the sample is marked in a residual variance versus Hotelling’s T 2 plot, and then the
model is recalculated without the selected sample.
While the residual variance versus Hotelling’s T 2 plot is very efficient in identifying obvious outliers, it is important to underline that there is no general method
for outlier recognition and removal. This is because, among other things, Hotelling’s
T 2 “outliers” may be desirable as extreme but valid specimens that span the model.
7.4.5 PCA for Data Quality Control
Due to its capability to model-free convey the samples inter-variability, PCA is a
very effective tool for quality control of an experimental dataset. Not only can PCA
be used to detect outliers as described above, but it also provides information on how
samples are related to each other in a quantitative series (such as in Fig. 7.17), in a
time series or in discrete groupings, which by PCA can all be scrutinized concerning
the smallest detail. Browsing through the PCA plots of a newly recorded dataset
can usually reveal more information about the data, than is otherwise possible from
Précédent

- 159/586

Suivant