162
K. M. Sørensen et al.
7.6 Validation of Multivariate Models
The purpose of validation of multivariate NIRS models is to provide an
unbiased evaluation of the model performance.
When selecting the validation method, you should act as the advocate of the devil!
A key concept in multivariate data analysis is model validation. Generally, validation is about the models’ applicability (extrapolation) to new samples. Unfortunately,
it is not enough to use the model fit alone as a validation criterion, since addition of
components will nearly always lead to an improved fit in the least squares sense for a
finite dataset. This means that we cannot use the calibration diagnostics to see if the
model is good or bad. A key question is instead “how will it perform for other and
new data?”. In order to try to answer this, it is necessary to find a way to validate the
number of components used in the multivariate model. If too few are used, the model
is said to be underfitted, and if too many, the model is said to be overfitted. The way
to estimate the correct number of components in a calibration model is to use a test
set, which is a set of sample spectra and related response variables unknown to the
model. The model can be developed on the calibration set, and the goodness can be
evaluated by the test set. This is called test set validation. Sometimes, when the total
dataset is not large enough to be split into both a calibration and test set, there exists
another option, which is called cross-validation. In this section, we will summarize
the most common types of validation employed in multivariate data analysis.
7.6.1 Model Performance Metrics
Correlation is a key statistic used to gauge regression model performance. Pearson’s
correlation coefficient, or just R, of known y and the associated predicted ˆ
y is:
R =
cov
y, ˆ
y
std(y) · std
ˆ
y
(7.26)
Sometimes expressed as coefficient of determination, R
2 provides a measure for
the relationship between the predicted outcome and the reference. R
2
= 1 corresponds
to a perfect relation, and R
2
= 0 corresponds to no relationship at all. Often, it is stated
as a stand-alone indicator for regression accuracy. However, as shown in Fig. 7.21,
it is a dangerous assumption to equate a high correlation to a high model quality. In
most practical circumstances, a R = 0.8 will be considered high, but as it is seen in
Fig. 7.21, it hardly reflects a metric that resembles quality.
Clearly, the correlation does indicate the accuracy of the given prediction. A
prediction may be near perfectly related to the associated reference values (high R),
K. M. Sørensen et al.
7.6 Validation of Multivariate Models
The purpose of validation of multivariate NIRS models is to provide an
unbiased evaluation of the model performance.
When selecting the validation method, you should act as the advocate of the devil!
A key concept in multivariate data analysis is model validation. Generally, validation is about the models’ applicability (extrapolation) to new samples. Unfortunately,
it is not enough to use the model fit alone as a validation criterion, since addition of
components will nearly always lead to an improved fit in the least squares sense for a
finite dataset. This means that we cannot use the calibration diagnostics to see if the
model is good or bad. A key question is instead “how will it perform for other and
new data?”. In order to try to answer this, it is necessary to find a way to validate the
number of components used in the multivariate model. If too few are used, the model
is said to be underfitted, and if too many, the model is said to be overfitted. The way
to estimate the correct number of components in a calibration model is to use a test
set, which is a set of sample spectra and related response variables unknown to the
model. The model can be developed on the calibration set, and the goodness can be
evaluated by the test set. This is called test set validation. Sometimes, when the total
dataset is not large enough to be split into both a calibration and test set, there exists
another option, which is called cross-validation. In this section, we will summarize
the most common types of validation employed in multivariate data analysis.
7.6.1 Model Performance Metrics
Correlation is a key statistic used to gauge regression model performance. Pearson’s
correlation coefficient, or just R, of known y and the associated predicted ˆ
y is:
R =
cov
y, ˆ
y
std(y) · std
ˆ
y
(7.26)
Sometimes expressed as coefficient of determination, R
2 provides a measure for
the relationship between the predicted outcome and the reference. R
2
= 1 corresponds
to a perfect relation, and R
2
= 0 corresponds to no relationship at all. Often, it is stated
as a stand-alone indicator for regression accuracy. However, as shown in Fig. 7.21,
it is a dangerous assumption to equate a high correlation to a high model quality. In
most practical circumstances, a R = 0.8 will be considered high, but as it is seen in
Fig. 7.21, it hardly reflects a metric that resembles quality.
Clearly, the correlation does indicate the accuracy of the given prediction. A
prediction may be near perfectly related to the associated reference values (high R),
