7 NIR Data Exploration and Regression by Chemometrics—A Primer
175
1100
1500
2000
2500
Wavelength [nm]
0
2
4
6
8
10
VIP Scores
2244 nm
1 2
4
6
8
10
Number of components
1.5
2
2.5
RMSE
RMSEC
RMSECV
0
2 0
4 0
6 0
8 0
%DE Measured
0
20
40
60
80
%DE CV predicted
a
b
c
RMSECV = 2.22
R 2 = 0.98
2 Components
Fig. 7.29 VIP score variable selection of the PLS model to the pectin %DE in the combined Dataset
1. a The model VIP scores, where the red dotted line at y = 4 (selected in this instance as half the
height of the highest VIP score) indicates the inclusion threshold. b The model statistics: RMSECV
of 2.22%DE with a two-component model is the optimal choice. c The resulting actual versus
predicted plot with an R2 of 0.98
variables with insignificant regression coefficients in order to improve the predictive
power of the model. However, as they are based on regression against y, degrees of
freedom are lost in the process and the produced models should be validated using
a test set to secure that the model is not overfitted.
Figure 7.29 shows the application of VIP score variable selection to Dataset 1.
The performance when imposing the variable selection (VIP scores > 4 which result
in 6 variables retained), the model has a similar RMSECV of 2.22%DE to the global
PLS model (Fig. 7.28) but uses one component less. This is a typical result of variable
selection; i.e., a few variables are selected (good for interpretation), the regression
model is deteriorated a bit (not good for scrutiny of performance but perhaps good
for robustness), and model is using fewer components (good for interpretation and
avoidance of interferences).
Variable selection based on VIP scores is becoming rather common and is available
in most commercial software.
7.7.3 Forward Stepwise Selection
The primary target of most variable selection methods is to improve the regression,
and this concept is employed in the most direct and brute way in the forward stepwise
selection (FSS) procedure. In this method, all single independent variables are tested
in finding the one, which provides the best regression model toward the dependent
variable. All these single-variable models are test set validated, and the variable with
the lowest RMSEP (on the independent test set) is chosen. In a second iteration,
one new variable, that improves the PLS model the best, is included to complement
the first selected variable. The variable that, in combination with the first chosen
variable, gives the lowest RMSEP is selected. Subsequently, additional variables
are included one-by-one by their capability to improve the previous model. This
175
1100
1500
2000
2500
Wavelength [nm]
0
2
4
6
8
10
VIP Scores
2244 nm
1 2
4
6
8
10
Number of components
1.5
2
2.5
RMSE
RMSEC
RMSECV
0
2 0
4 0
6 0
8 0
%DE Measured
0
20
40
60
80
%DE CV predicted
a
b
c
RMSECV = 2.22
R 2 = 0.98
2 Components
Fig. 7.29 VIP score variable selection of the PLS model to the pectin %DE in the combined Dataset
1. a The model VIP scores, where the red dotted line at y = 4 (selected in this instance as half the
height of the highest VIP score) indicates the inclusion threshold. b The model statistics: RMSECV
of 2.22%DE with a two-component model is the optimal choice. c The resulting actual versus
predicted plot with an R2 of 0.98
variables with insignificant regression coefficients in order to improve the predictive
power of the model. However, as they are based on regression against y, degrees of
freedom are lost in the process and the produced models should be validated using
a test set to secure that the model is not overfitted.
Figure 7.29 shows the application of VIP score variable selection to Dataset 1.
The performance when imposing the variable selection (VIP scores > 4 which result
in 6 variables retained), the model has a similar RMSECV of 2.22%DE to the global
PLS model (Fig. 7.28) but uses one component less. This is a typical result of variable
selection; i.e., a few variables are selected (good for interpretation), the regression
model is deteriorated a bit (not good for scrutiny of performance but perhaps good
for robustness), and model is using fewer components (good for interpretation and
avoidance of interferences).
Variable selection based on VIP scores is becoming rather common and is available
in most commercial software.
7.7.3 Forward Stepwise Selection
The primary target of most variable selection methods is to improve the regression,
and this concept is employed in the most direct and brute way in the forward stepwise
selection (FSS) procedure. In this method, all single independent variables are tested
in finding the one, which provides the best regression model toward the dependent
variable. All these single-variable models are test set validated, and the variable with
the lowest RMSEP (on the independent test set) is chosen. In a second iteration,
one new variable, that improves the PLS model the best, is included to complement
the first selected variable. The variable that, in combination with the first chosen
variable, gives the lowest RMSEP is selected. Subsequently, additional variables
are included one-by-one by their capability to improve the previous model. This
