166
K. M. Sørensen et al.
1
2
4
6
8
1 0
Components in model
0.12
0.14
0.16
0.18
0.2
0.22
0.24
0.26
0.28
0.3
0.32
RMSEC, RMSECV
RMSEC
RMSECV
Fig. 7.23 Model fit as a function of number of components. A typical development of the root mean
square error of calibration (RMSEC) and root mean square error of cross-validation as a function of
the number of PLS components. The RMSECV line has a local minimum at component #3, which
is optimal for the model
7.6.4 Cross-Validation Systems
The strength of cross-validation lies in testing the prediction on samples independent
from the calibration. It requires that the cross-validation segments are constructed in
proper correspondence with the experiment—especially if the experiment includes
replicate measurements. If 20 samples are measured with five replicates, a regression
on the 100 actual sample measurements should be cross-validated in a way so that all
five replicates of each sample are removed as a group (20 cross-validation segments),
and not by removing each 20th object as a group (5 cross-validation segments). Only
the former setup will give independent evaluation of each sample, which is not the
case in the latter case, where 4 versions of the same sample are still included per
CV segment, which then no longer can be called independent anymore. The former
case will test the model’s ability to predict unknown samples, where the latter will
validate the model against the measurement and tolerances, as this is the changing
factor between the cross-validation segments.
Several systems specifying the segmentation of the samples in cross-validation
setups are used in the literature and software. These include venetian blinds, where
the CV segments are grouped in sets 1-2-3-1-2-3-1-2-3-…, contiguous subsets where
the CV segments are grouped in sets 1-1-1-2-2-2-3-3-3-…, or combinations thereof,
K. M. Sørensen et al.
1
2
4
6
8
1 0
Components in model
0.12
0.14
0.16
0.18
0.2
0.22
0.24
0.26
0.28
0.3
0.32
RMSEC, RMSECV
RMSEC
RMSECV
Fig. 7.23 Model fit as a function of number of components. A typical development of the root mean
square error of calibration (RMSEC) and root mean square error of cross-validation as a function of
the number of PLS components. The RMSECV line has a local minimum at component #3, which
is optimal for the model
7.6.4 Cross-Validation Systems
The strength of cross-validation lies in testing the prediction on samples independent
from the calibration. It requires that the cross-validation segments are constructed in
proper correspondence with the experiment—especially if the experiment includes
replicate measurements. If 20 samples are measured with five replicates, a regression
on the 100 actual sample measurements should be cross-validated in a way so that all
five replicates of each sample are removed as a group (20 cross-validation segments),
and not by removing each 20th object as a group (5 cross-validation segments). Only
the former setup will give independent evaluation of each sample, which is not the
case in the latter case, where 4 versions of the same sample are still included per
CV segment, which then no longer can be called independent anymore. The former
case will test the model’s ability to predict unknown samples, where the latter will
validate the model against the measurement and tolerances, as this is the changing
factor between the cross-validation segments.
Several systems specifying the segmentation of the samples in cross-validation
setups are used in the literature and software. These include venetian blinds, where
the CV segments are grouped in sets 1-2-3-1-2-3-1-2-3-…, contiguous subsets where
the CV segments are grouped in sets 1-1-1-2-2-2-3-3-3-…, or combinations thereof,
