7 NIR Data Exploration and Regression by Chemometrics—A Primer
129
analysis of NIR spectra [2–4]. While many of these are indeed excellent, many
newcomers are understandably overwhelmed by the amount of equations and abundance of methodologies, guidelines and recommendations. To assist beginners in
multivariate analysis of NIR data, this chapter instead takes the form of a pragmatic
step-by-step tutorial for multivariate analysis of ensembles of NIR spectra obtained
on similar classes of samples.
In classical empirical research, a model requires that the number of variables
must be less than or equal to the number of observations. In spectroscopy, this would
correspond to the number of intensities measured at individual wavelengths and
number of samples. If these samples are measured by NIR spectroscopy, such as
in a conventional quality control analysis setup, at least 1000 spectral variables are
recorded. A typical dataset of 100 samples will thus have the dimensions 100 ×
1000, incompatible with traditional empirical models, but effectively dealt with using
chemometrics, which can handle collinear data structures with many more variables
than samples.
The trick in chemometrics is the reduction of the complex dataset into a limited set
of latent (or principal) variables, which in turn can be used to model (unsupervised)
and visualize class belonging, identify outliers, suggest trends, etc. If response variables measured by another reference method are available as well, the latent structures
can be used for supervised regression modeling and prediction. In general, chemometrics assumes additivity of underlying components (in spectroscopic terms called
Lambert–Beer law) and bilinear relations between the spectra (X) and response
variables (y). Within this framework of the “straight-line tyranny,” we generally
decompose the spectral datasets as follows:
X = A · B
T
+ E
(7.1)
where X (n x m) is a set of n sample spectra x of length m:
X =
⎡
⎢
⎣
x 1,1 · · · x 1,m
. . .
. . .
. . .
x n,1 · · · x n,m
⎤
⎥
⎦
(7.2)
and A (n x f ) and B (m x f ) contain f latent variables (see Fig. 7.2).
A, with typically f « n and m, then contains the contributions (or pseudoconcentrations) of the hidden latent phenomenon modeled by B. E contains the residuals or unexplained information (in a least squares sense). Appearing some 40 years
ago on the scientific scene [5], chemometrics has established itself as an indisputable
effective and valuable multivariate data analysis toolset, extensively used in spectroscopic applications to extract information from data that would otherwise remain
hidden for classical univariate methods. It exploits the multivariate advantage, while
at the same time facilitating noise reduction and allowing for outlier removal.
Before going into deeper facets of chemometrics, we will briefly introduce the
datasets that will be used throughout this chapter to illustrate the different algorithms.
Précédent

- 134/586

Suivant