174
K. M. Sørensen et al.
High absolute regression coefficients are considered important, and small regression
coefficients are considered less important and could potentially be eliminated. Eliminating variables that have a small absolute value for the regression coefficients
may thus lead to an improvement of the model [46]. This simple idea is the basis
for variable selection by so-called jackknifing [47, 48]. For each cross-validation
segment, a new regression vector is calculated, and it will thus be possible to estimate the standard deviation of the PLS regression vector. When the distribution of
the regression coefficients includes zero, they can be discarded, and a new model
calculated. The procedure can be implemented iteratively by recalculating the model
and eliminating more variables. It differs from other variable selection methods by
not searching directly for variables that are good at predictions or y, but instead by
eliminating variables that possibly have a regression coefficient close to zero and thus
do not contribute (significantly) to the prediction. No matter what value such variables have for a new sample, they will be multiplied by zero and thus not contribute
in the prediction. To remove variables, this way makes the risk of overfitting smaller
compared to selection methods that are focused on finding variables that are good at
predicting the response variable.
It is a prerequisite for this variable selection method to work, that a decent model
has already been developed using all spectral variables. As shown in Fig. 7.28c, the
majority of the regression coefficients are close to zero despite the fact that the raw
spectral data varying significantly in these regions. This indicates that the PLS model
has down-weighted these variables.
Since spectroscopic data are normally smooth, so is the regression vector expected
to appear smooth. It is thus (normally) not possible that the measurements at, e.g.,
2450 nm are positively correlated to the response variable and the measurements at
2452 nm are negatively correlated. Models with noisy regression coefficients indicate
that the spectral region should be removed. By these very pragmatic and basic rules,
it is possible to reduce the spectral data to a region of interest, which improve the
predictive PLS model. This method further has the practical advantage that PLS
regression coefficients are readily available as standard output and plots from PLSR
software.
7.7.2 Variable Importance in Projection
A variant in using the regression coefficients for variable selection is the variable
importance in projection (VIP) estimate, originally proposed by Wold et al. [49].
The VIP score is an estimate of the importance of each variable for the projection of
y onto X. It is found by accumulating the importance of each variable from the PLS
loading weights for each component. The average VIP score for all variables is equal
to 1, and hence, typically a “larger-than-one” selection rule is applied for variable
selection [50]. The VIP score is normally used as an assistance in manual variable
selection and can be a valuable tool, when used together with prior knowledge about
the measurements. As a selectivity ratio, the VIP number can be used to exclude
K. M. Sørensen et al.
High absolute regression coefficients are considered important, and small regression
coefficients are considered less important and could potentially be eliminated. Eliminating variables that have a small absolute value for the regression coefficients
may thus lead to an improvement of the model [46]. This simple idea is the basis
for variable selection by so-called jackknifing [47, 48]. For each cross-validation
segment, a new regression vector is calculated, and it will thus be possible to estimate the standard deviation of the PLS regression vector. When the distribution of
the regression coefficients includes zero, they can be discarded, and a new model
calculated. The procedure can be implemented iteratively by recalculating the model
and eliminating more variables. It differs from other variable selection methods by
not searching directly for variables that are good at predictions or y, but instead by
eliminating variables that possibly have a regression coefficient close to zero and thus
do not contribute (significantly) to the prediction. No matter what value such variables have for a new sample, they will be multiplied by zero and thus not contribute
in the prediction. To remove variables, this way makes the risk of overfitting smaller
compared to selection methods that are focused on finding variables that are good at
predicting the response variable.
It is a prerequisite for this variable selection method to work, that a decent model
has already been developed using all spectral variables. As shown in Fig. 7.28c, the
majority of the regression coefficients are close to zero despite the fact that the raw
spectral data varying significantly in these regions. This indicates that the PLS model
has down-weighted these variables.
Since spectroscopic data are normally smooth, so is the regression vector expected
to appear smooth. It is thus (normally) not possible that the measurements at, e.g.,
2450 nm are positively correlated to the response variable and the measurements at
2452 nm are negatively correlated. Models with noisy regression coefficients indicate
that the spectral region should be removed. By these very pragmatic and basic rules,
it is possible to reduce the spectral data to a region of interest, which improve the
predictive PLS model. This method further has the practical advantage that PLS
regression coefficients are readily available as standard output and plots from PLSR
software.
7.7.2 Variable Importance in Projection
A variant in using the regression coefficients for variable selection is the variable
importance in projection (VIP) estimate, originally proposed by Wold et al. [49].
The VIP score is an estimate of the importance of each variable for the projection of
y onto X. It is found by accumulating the importance of each variable from the PLS
loading weights for each component. The average VIP score for all variables is equal
to 1, and hence, typically a “larger-than-one” selection rule is applied for variable
selection [50]. The VIP score is normally used as an assistance in manual variable
selection and can be a valuable tool, when used together with prior knowledge about
the measurements. As a selectivity ratio, the VIP number can be used to exclude
