164
K. M. Sørensen et al.
7.6.2 Model Validation
In order to develop reliable error measurements for a regression, the model must
undergo a validation, where it is tested how it predicts unknown or new data. In this
respect, new or unknown refers to data points which were not included in the dataset
used while calibrating the model. That does not exclude them from originating from
the same experiment. However, a very crucial rule needs to be enforced for any split
of data into a calibration and validation set: The validation data must be completely
independent from the calibration data. In practice, this means that the same physical
sample measured cannot be present in both datasets, nor as separate measurements
and neither as measurement replicates. A good way of thinking of this is that the data
conceptually could come from two different measurement campaigns.
As described below, one does not always have the luxury of an isolated dataset
used for validation purposes. In those cases, the method of cross-validation can be
applied to get validation statistics. But, however way the data are validated, it is
important to stress that the chosen validation scheme is determined by the structure
of the experiment so as to ensure independent validation data, and hence should be
a consideration made very early in any calibration workflow.
A term frequently used in multivariate analysis is the concept of overfitting. A
model is said to be overfitted, when it loses the ability to optimally predict new
or unknown samples. This is particularly relevant to regression models, where the
inclusion of too many components will cause the regression to overfit the data and
essentially making it useless for any practical purpose. When the model overfits
the data, the additional components will start to structure, not just the systematic
variation in the data, but also the sample specific noise.
7.6.3 Cross-Validation
In order to avoid overfitting, a model can be validated in one of two ways, namely by
applying it on a test set or by cross-validation. Cross-validation (CV) is the process
of sequentially removing one or more samples, makes a calibration on the remaining
samples and uses that to predict the values of the ones removed [42, 43]. The process
is illustrated in Fig. 7.22.
Cross-validation can be used to find an optimal number of components, when the
RMSE of cross-validated predicted values (named RMSECV) is calculated against
the associated reference (original y). When RMSECV is plotted against the number of
components included in the model, it will typically reveal a local minimum that yields
the most accurate model. A characteristic RMSECV development, along with its
corresponding root mean square error of calibration (RMSEC), is shown in Fig. 7.23.
As may be expected, the error occurring from the unvalidated calibration (blue line)
Précédent

- 169/586

Suivant