amount of explained variance). Blanchet et al. (2008a) addressed this double problem and proposed solutions to improve this technique:
• To prevent the problem of inflation of the overall type I error, a global test using
all explanatory variables is first run. If, and only if, that test is significant, the
forward selection is performed;
• To reduce the risk of incorporating too many variables into the model, the
adjusted coefficient of multiple determination R
2
adj of the global model
(containing all the potential explanatory variables) is computed, and used as a
second stopping criterion. Forward selection is stopped if one of the following
criteria is reached: the α significance level or the global R
2
adj ; in other words, if a
candidate variable is deemed nonsignificant or if it brings the R
2
adj of the current
model over the value of the R
2
adj of the global model.
Another criterion applied in model selection is Akaike’s information criterion
(AIC). In multivariate applications like RDA, an AIC-like criterion can be
computed, but, according to Oksanen (2017, documentation file of function
ordistep()in vegan) it may not be completely trustworthy. Furthermore,
experience shows that it tends to be very liberal.
Three functions are mostly used in ecology for variable selection in the RDA
context: forward.sel()of package adespatial and ordistep()and its
offspring ordiR2step()of package vegan. Let us explore these functions by
presenting them and applying them each in turn.
Forward selection with function forward.sel()
Function forward.sel()requires a response data matrix and an explanatory data
matrix which, unfortunately, must contain quantitative variables only. Factors must
be recoded in the form of dummy variables.
To identify the “best” explanatory variables in turn, and decide when to stop the
selection, forward.sel()finds the explanatory variable with the highest R
2 (first
variable) or partial (additional) R
2 (for the following variables). The decision to
retain a variable or to stop the procedure can be made on several criteria: a
pre-selected significance level α (argument alpha), and the adjusted coefficient
of multiple determination R
2
adj of the global model containing all the potential
explanatory variables (argument adjR2thresh). Note that forward.sel()
allows for additional, more arbitrary stopping criteria such as a maximum number of
variables entered, a maximum R
2 and a minimum additional contribution to the R
2 .
These criteria are rarely used.
Let us apply forward.sel() to our fish and environmental data. Since this
procedure does not allow for factor variables, we will use the env2 data set, which
contains only quantitative variables, in the analysis.
6.3 Redundancy Analysis (RDA)
227
• To prevent the problem of inflation of the overall type I error, a global test using
all explanatory variables is first run. If, and only if, that test is significant, the
forward selection is performed;
• To reduce the risk of incorporating too many variables into the model, the
adjusted coefficient of multiple determination R
2
adj of the global model
(containing all the potential explanatory variables) is computed, and used as a
second stopping criterion. Forward selection is stopped if one of the following
criteria is reached: the α significance level or the global R
2
adj ; in other words, if a
candidate variable is deemed nonsignificant or if it brings the R
2
adj of the current
model over the value of the R
2
adj of the global model.
Another criterion applied in model selection is Akaike’s information criterion
(AIC). In multivariate applications like RDA, an AIC-like criterion can be
computed, but, according to Oksanen (2017, documentation file of function
ordistep()in vegan) it may not be completely trustworthy. Furthermore,
experience shows that it tends to be very liberal.
Three functions are mostly used in ecology for variable selection in the RDA
context: forward.sel()of package adespatial and ordistep()and its
offspring ordiR2step()of package vegan. Let us explore these functions by
presenting them and applying them each in turn.
Forward selection with function forward.sel()
Function forward.sel()requires a response data matrix and an explanatory data
matrix which, unfortunately, must contain quantitative variables only. Factors must
be recoded in the form of dummy variables.
To identify the “best” explanatory variables in turn, and decide when to stop the
selection, forward.sel()finds the explanatory variable with the highest R
2 (first
variable) or partial (additional) R
2 (for the following variables). The decision to
retain a variable or to stop the procedure can be made on several criteria: a
pre-selected significance level α (argument alpha), and the adjusted coefficient
of multiple determination R
2
adj of the global model containing all the potential
explanatory variables (argument adjR2thresh). Note that forward.sel()
allows for additional, more arbitrary stopping criteria such as a maximum number of
variables entered, a maximum R
2 and a minimum additional contribution to the R
2 .
These criteria are rarely used.
Let us apply forward.sel() to our fish and environmental data. Since this
procedure does not allow for factor variables, we will use the env2 data set, which
contains only quantitative variables, in the analysis.
6.3 Redundancy Analysis (RDA)
227
