7 NIR Data Exploration and Regression by Chemometrics—A Primer
167
as 1-1-2-2-3-3, etc. This nomenclature refers to how the data are structured in the
matrices and must be accompanied by other information (number of splits, number
of segments, etc.). Metadata can often be used to split the data into sensible nested
segments (experimental design factors) such as e.g. animals, varieties, vintages
etc. Full cross-validation (or leave-one-out/LOO cross-validation) is often referenced
in the literature, but should be used with caution, especially in the case of data with
analytical replicates. Leaving out every single sample does rarely provide independent sampling, and the full cross-validation should only be applied in cases with very
few samples [44].
7.6.5 Bootstrapping
An alternative to full cross-validation, or when no prior knowledge of the data structure is available, is to divide the dataset into a number of random blocks of each
typically 10–20% of the data, called random subsets. In case of replicates, all replicates of the same sample must still be kept out at the same time. Repeating the
random sampling validation, a high number of times for a dataset, each time with
new randomization, gives a robust error estimate by averaging the CV predicted y
over the repeated CV runs.
7.6.6 Test Set Validation
When enough samples are available, or when the experiment design permits, a very
efficient way of validating a model is to do test set validation. Ideally, the experimental
data can be split into independent calibration and test parts, each representative of
the population of observations in the experiment. As described earlier, replicate
measurements of the same physical sample cannot be present in both sets at the same
time.
Test set validation is straightforward. A model is calculated for the data in the
calibration dataset, which is then applied to the test set. The error of the predicted
test set values and the associated reference values determined is referred to as root
mean square error of prediction (RMSEP).
7.6.7 Application of PLS to NIR Spectra
In a first example, the application of PLS to the mixture design in Dataset 2 is
demonstrated, using the NIR spectra as X and the glucose content as the response
variable, y. The results of this model are shown in Fig. 7.24.
Précédent

- 172/586

Suivant