7 NIR Data Exploration and Regression by Chemometrics—A Primer
161
When you ask for discrimination, you will get it!
—Lars Nørgaard, Danish chemometrician
While PCA results often can be presented without considering validation, PLS
models must always be validated before presenting scores and loadings and prediction
errors (see section on validation). For PLS-DA models, the validation becomes even
more crucial since spurious correlations can often lead to excellent, but false, classifications. Moreover PLS-DA score plots should be used with great care (read: not be
used) since it can be demonstrated that score plots from a PLS-DA model often can
show clear groupings even when random data is assigned to two classes. Similarly,
discriminative PLS-DA score plot can be found when sound real data are arbitrarily
divided into two classes [38]. Regardless of validation or not, the scores and loading
plots would be similar and these plots can thus not be used to access the classification
performance of a PLS-DA model.
7.5.4 Outro
In many practical applications, multiple response variables are available and for this
purpose there is a variant of PLS called PLS2, which can be used as alternative. It
could, for example, be that one would like to predict protein, fat and carbohydrate
content of a cereal product. With the help of PLS2, these three different models can
be made at once and thus used to directly understand how the three different quality
parameters interact. However, if performance is the single objective, then it is highly
likely that you will get better performance results by just applying PLS separately to
each of the three response variables.
For many PLS applications, the target is to minimize the prediction error. It should
be as low as possible, yet maintaining its predictive power. It is important to note that
PLS is correlation/covariance-based and will not be able to distinguish between direct
correlations (causal) and indirect correlations. A sound and healthy PLS regression
model may very well rely on indirect (biological) correlations in the sample set—also
called the cage of covariance [39].
PLS is implemented in all chemometric software and is probably the strongest
regression tool ever developed. Accordingly, good reasons (typically called nonlinearities) are needed for not choosing PLS in multivariate regression. Other alternatives are principal component regression, random forest, neural networks and
machine learning, which will be briefly discussed later in this chapter.
Précédent

- 166/586

Suivant