VIFs can also be computed in vegan after RDA or CCA. The algorithm in the
vif.cca() function allows users to include factors in the RDA; the function will
compute VIF after breaking down each factor into dummy variables. If X contains
quantitative variables only, the vif.cca() function produces the same results as
the equation above. Here is an example where VIFs are computed after RDA for a
matrix X containing quantitative variables plus a factor (the variable “slo”). Several
VIF values are above 10 or even 20 in these analyses, so that a reduction of the
number of explanatory variables is justified.
# Variance inflation factors (VIF) in two RDAs
# First RDA of this Chapter: all environmental variables
# except dfs
vif.cca(spe.rda)
# Partial RDA – physiographic variables only
vif.cca(spechem.physio)
No single, perfect method exists to reduce the number of variables, besides the
examination of all possible subsets of explanatory variables, which is timeprohibitive for real data sets. In multiple regression, the three usual methods are
forward selection, backward elimination and stepwise selection of explanatory variables, the latter being a combination of the first two. In RDA, forward selection is the
method most often applied because it works even in cases where the number of
explanatory variables is larger than (n – 1). This method works as follows:
• Compute m RDAs of the response data with one of the m explanatory variables
in turn.
• Select the “best” explanatory variable on the basis of criteria that are developed
below. If it is significant. . .
• . . . the next task is to look for a 2nd (3rd, 4th, etc.) variable to include in the
explanatory model. Compute all models containing the previously selected variable(s) plus one of the remaining explanatory variables. Identify the “best” new
variable; include it in the model if its partial contribution is significant at the
pre-selected significance level.
• The process continues until no more significant variable can enter the model.
Several criteria exist to decide when to stop variable selection. The traditional one
is a pre-selected significance level α (selection is stopped when no additional
variable has a partial permutational p-value smaller than or equal to α). However,
this criterion is known to be overly liberal, either by sometimes selecting a “significant” model when none should have been identified (hence inflating type I error), or
by including too many explanatory variables into the model (hence inflating the
226
6 Canonical Ordination
Précédent

- 238/444

Suivant