180
K. M. Sørensen et al.
1. To improve the performance of multivariate regression models
(RMSEP/RMSECV)
2. To simplify multivariate regression models by excluding interferences, local
nonlinearities and noisy variables (fewer latent variables)
3. To improve interpretability of multivariate models.
It is normally not a good idea to compare and scrutinize variable selection methods
for better performance. The danger of overfitting is too high, and the applicability of
the final models will often become too limited and sensitive. The pragmatic compromise is often to use a variation of iPLS in which a spectral region can be selected
with the relevant signals and without deteriorating noise or interferences present.
Remember the instrumental spectral range was not decided for a specific application!
However, when seeking for causality and interpretation, variable selection methods
may be a strong tool to combine with a priori knowledge.
7.8 ANOVA Simultaneous Component Analysis (ASCA)
While Variable Selection can be considered as a horizontal elimination
of interferences, ASCA can be considered as a vertical elimination of
interferences (partition of variances)
As mentioned previously, the most valuable meta-parameters in any spectral
recording sets are the experimental design factors. It is a good practice to use this
knowledge, and many NIRS studies provide multivariate datasets with an underlying
experimental design.
Biological systems exhibit sources of variation due to a large number of factors,
such as variety, soil and climate. Realizing this, led Fisher [53] (broadly recognized
as the father of modern statistics), to develop experimental designs suited for estimation and handling of the variation based on these factors. The paired t-test and
analysis of variance (ANOVA) are examples of models used to analyze data from
designed experiments. The backbone of these methods is to estimate variance related
to the design factors, including both systematic factors such as treatment, but also
a nuisance factor like subject, and hence in turn be able to remove dominating but
un-interesting variation. In this way, the variation of interest, like treatment, is emphasized, which in turn increases the chance of finding something interesting. For a wide
range of applications, the dominating variation in data is often trivial, while the interesting—and new—variation sources can be minor in comparison, leaving it covered
or unresolved if not handled through a proper designed experimental structure and
in turn elucidated via a mathematical extraction of the relevant design effects.
Précédent

- 185/586

Suivant