7 NIR Data Exploration and Regression by Chemometrics—A Primer
181
An interesting multivariate tool for exploiting the experimental design information, combining the power of ANOVA to separate variance sources with the advantages of simultaneous component analysis (SCA) to modeling of the individual separate effect matrices, is called ANOVA simultaneous component analysis (ASCA)
[54]. It utilizes advantages of ANOVA in terms of both partitioning the sources of
variance and using PCA for multivariate interpretation. In order to exemplify how
this method works, imagine a simple experiment, where n samples are treated by
process A and process B in a randomized crossover fashion. At the end of each
process, a quality is measured. If this response is univariate, the difference between
processes A and B is naturally tested by a t-test. The power of a t-test is that each
sample serves as its own control and the variation is hence split into what origins
from the individual samples and what origins from the processes. More formally, the
model can be written as:
x i = a
process i
+ β
sample i
+ e i
(7.30)
where α has k levels (the number of different processes) and β has j levels (the number
of samples). e i is the error term, which often is assumed normally distributed and
may have an arbitrary number of levels. In a testing scenario, the aim is to compare
the effect, i.e., differences between different levels of α, with the magnitude of the
(random) error.
If the response is multivariate (e.g., NIR spectrum), this paradigm simply scales to
the multivariate case. Take the setup from above, but let the response be multivariate
(X); the model of X can now be formalized as follows:
X = X(process) + X(sample) + E
(7.31)
For a full crossover, the dimension of the X’s is k j by m (m is the number of
variables). X(process) describes the information related to the different processes
and has k levels, while X(sample) describes between sample variations (j levels).
Equation 7.31 hence is merely a concatenation of Eq. 7.30 m times, one for each
variable. E represents the non-design-related information—it is often systematic but
just not related to the experimental design. The right-hand side of Eq. 7.31 can be
combined or analyzed individually by, for example, PCA (the working principle
in ASCA). Assume that X is made of NIR spectra, then a PCA on X process will
point toward spectral patterns that discriminate between the processes, and likewise
a PCA on X sample will reflect where the largest variation due to sample differences
(e.g., variety, soil, climate, etc.) is distributed. If the aim is to investigate the processrelated patterns by taking the error spread into account, a score plot made from
projecting X process + E onto loadings from a PCA on only X process will reflect the
process differences in relation to the non-design-related variation in the data. If
the aim is to test for differences, multivariate classification models can be built on
relevant parts of Eq. 7.31. For example, a PLS-DA on X proces + E for classification of
processes would point toward how strong the process-related variation is compared to
Précédent

- 186/586

Suivant