7 NIR Data Exploration and Regression by Chemometrics—A Primer
165
1
2
3
1
2
3
X
y
2
3
2
3
1
3
1
3
1
2
1
2
1
1
2
2
3
3
Calibration
data
Data included
in CV
submodels
CV Prediction
Data excluded
from CV
submodels
Fig. 7.22 Schematic overview of the cross-validation process. A dataset is split into blocks—here
three, which each is removed once from the dataset in turn. As each block is taken out, a model can
be developed on the two remaining blocks. The new sub-model can be used to predict the values of
the excluded block (CV prediction). After excluding all blocks in turn, a complete y vector of CV
predictions has been produced
keeps falling, as more and more components are included in the model. The crossvalidated model (red line) instead shows a characteristic low point after four components. From there on, the prediction error starts to increase, showing that the model is
overfitting the data. As a rule of thumb, one must select as few components in a model
as possible in order to eliminate the possibility of overfitting. The developing model
in Fig. 7.23 show clearly that the 3-component model is optimal. Including further
components will cause the RMSECV to increase and, hence, overfit. However, other
PLS models sometimes display an insignificant improvement in RMSECV when
going from a 3-component system to a 4-component model. It can then be argued
that the proper conservative choice is to use the more parsimonious 3-component
model. Despite being offered as an option in several software packages, automatic
selection of the number of components should only be considered as a guidance. An
automatic system can never replace prior knowledge of the samples, the measurement
system or reference values.
Précédent

- 170/586

Suivant