7 NIR Data Exploration and Regression by Chemometrics—A Primer
177
7.7.4 Recursively Weighted PLS (rPLS)
Recursively weighted PLS (rPLS) [51] combines the selection of variables by regression coefficients with an automatic iterative variable selection procedure. rPLS iteratively eliminates variables by using the regression coefficients to magnify important variables and thus down-weight less important variables. rPLS is based on a
process of repeated PLS models where the current regression coefficients are used
as cumulative weights on X:
X i = X i−1 · diag(b i−1 )
(7.29)
where X i is the weighted X and b i is the regression coefficient for iteration i. Using
this method, a reduced subset of variables is identified for regression by recursively reweighing of the independent variables (X) with the estimated regression coefficients
(b). The algorithm is started with a standard PLS model between X 1 (equal to X)
and y, giving b 1 . The re-weighting is recursively repeated until no further change in
the regression coefficients occurs. The result is a regression vector b end that contains
only ones and zeros. This binary result is a direct output from the rPLS algorithm;
i.e., no rescaling of the final regression vector is performed. The rPLS model has
the advantage that it, under normal conditions, will converge to a limited number of
variables, normally including colinear neighbor variables, which is very useful for
interpretation. This is not the case in more complicated situations.
The method will ultimately and normally converge to a solution that has the same
number of variables as the number of principal components included in the regression.
rPLS has the advantage in comparison with other iterative variable selection methods
that no meta-parameters are required (i.e., interval sizes or number of components),
at the “optimum” a relative low number of variables will be included in the model,
and after recursive convergence very few variables are retained in the end model.
In the latter case, the model performance is slightly worse than the optimal, but the
interpretability may be significantly improved.
Figure 7.31 shows the application of rPLS variable selection to Dataset 1 using
only 2 components. The optimal performance is reached for iteration #7 and gives
a performance of RMSECV of 1.77%DE and a R
2 of 0.97 which is markedly better
than the global PLS model (Fig. 7.28) using only 2 components and 13 variables
(centered around 1460 nm and 2244 nm). Then, the method is allowed to converge,
and it reaches 3 variables (1460 and 2244 nm) and a performance of 1.78%DE. The
model results thus give excellent interpretation and demonstrate that the multivariate
advantage may be gained by just adding a few or a single covariate neighbor variable.
Again, the improvement in performance should give serious concern to the danger of
overfitting and the model will need a real test set to be confirmed. The model shown
in this section was determined in MATLAB using the open-source rPLS algorithm
available at http://www.models.life.ku.dk/algorithms.
177
7.7.4 Recursively Weighted PLS (rPLS)
Recursively weighted PLS (rPLS) [51] combines the selection of variables by regression coefficients with an automatic iterative variable selection procedure. rPLS iteratively eliminates variables by using the regression coefficients to magnify important variables and thus down-weight less important variables. rPLS is based on a
process of repeated PLS models where the current regression coefficients are used
as cumulative weights on X:
X i = X i−1 · diag(b i−1 )
(7.29)
where X i is the weighted X and b i is the regression coefficient for iteration i. Using
this method, a reduced subset of variables is identified for regression by recursively reweighing of the independent variables (X) with the estimated regression coefficients
(b). The algorithm is started with a standard PLS model between X 1 (equal to X)
and y, giving b 1 . The re-weighting is recursively repeated until no further change in
the regression coefficients occurs. The result is a regression vector b end that contains
only ones and zeros. This binary result is a direct output from the rPLS algorithm;
i.e., no rescaling of the final regression vector is performed. The rPLS model has
the advantage that it, under normal conditions, will converge to a limited number of
variables, normally including colinear neighbor variables, which is very useful for
interpretation. This is not the case in more complicated situations.
The method will ultimately and normally converge to a solution that has the same
number of variables as the number of principal components included in the regression.
rPLS has the advantage in comparison with other iterative variable selection methods
that no meta-parameters are required (i.e., interval sizes or number of components),
at the “optimum” a relative low number of variables will be included in the model,
and after recursive convergence very few variables are retained in the end model.
In the latter case, the model performance is slightly worse than the optimal, but the
interpretability may be significantly improved.
Figure 7.31 shows the application of rPLS variable selection to Dataset 1 using
only 2 components. The optimal performance is reached for iteration #7 and gives
a performance of RMSECV of 1.77%DE and a R
2 of 0.97 which is markedly better
than the global PLS model (Fig. 7.28) using only 2 components and 13 variables
(centered around 1460 nm and 2244 nm). Then, the method is allowed to converge,
and it reaches 3 variables (1460 and 2244 nm) and a performance of 1.78%DE. The
model results thus give excellent interpretation and demonstrate that the multivariate
advantage may be gained by just adding a few or a single covariate neighbor variable.
Again, the improvement in performance should give serious concern to the danger of
overfitting and the model will need a real test set to be confirmed. The model shown
in this section was determined in MATLAB using the open-source rPLS algorithm
available at http://www.models.life.ku.dk/algorithms.
