44
Jane Elith
Predictor Variables
There is a range of approaches to selecting predictor variables. Variable selection
is important when large amounts of environmental data are available but not
necessarily relevant. GLMs and GAMs are the only methods with widely recognized approaches to selecting subsets of variables—the most common of these,
forms of stepwise selection, are often available as options within the software
used for modeling. For the other methods, there are no inbuilt procedures for
variable selection. Nevertheless, there are several approaches including expert
judgment, principal components analysis, and univariate regression that could be
used to select variables prior to modeling with these methods.
GLMs and GAMs are the only methods with recognized practices for investigating interactions between variables. Several of the climate parameters in
ANUCLIM combine data in an interaction-like relationship (e.g., mean temperature of the wettest quarter) and clearly new variables that represent interactions
can be created for any of the methods, but there is no inbuilt capability for such
analysis.
There are a variety of approaches to working with correlated (collinear) predictor variables. ANUCLIM recommends the use of all climate parameters even
though the majority of them are often highly correlated. DOMAIN and GARP
have no recommended approach. Most texts on regression advise that dependencies among the covariates should be investigated (see, e.g., Hosmer and
Lemeshow 1989; Agresti 1996), but many statistical packages do not have inbuilt
routines for correlation analysis that are part of the modeling component. This
leaves it to the practitioners to develop their own good practice (see, e.g., Booth et
al. 1994).
Goodness of Fit
GLMs and GAMs are the only methods that provide statistical estimates of error
and deviance. These report the extent to which the model does not fit the data.
Usability and Transparency
All methods require some computing and statistical understanding. All can be
mastered with some effort, and none requires expert programming knowledge.
The time required to achieve proficiency will clearly vary with the user’s experience. The more important consideration is that the method and its output are well
understood by the user. Understanding is affected by the complexity of the underlying concepts, the transparency of the procedures, the form of the final model, the
extent of documentation of the particular programs used, the number of alternative packages for implementing the method (can the results be replicated?), and
the breadth of published applications of the method. Thus potential difficulties
with the more complex statistical concepts of the GLMs and GAMs are balanced
by the extensive documentation and literature associated with them (although
GAMs are much newer and are less commonly implemented), and the user-
Jane Elith
Predictor Variables
There is a range of approaches to selecting predictor variables. Variable selection
is important when large amounts of environmental data are available but not
necessarily relevant. GLMs and GAMs are the only methods with widely recognized approaches to selecting subsets of variables—the most common of these,
forms of stepwise selection, are often available as options within the software
used for modeling. For the other methods, there are no inbuilt procedures for
variable selection. Nevertheless, there are several approaches including expert
judgment, principal components analysis, and univariate regression that could be
used to select variables prior to modeling with these methods.
GLMs and GAMs are the only methods with recognized practices for investigating interactions between variables. Several of the climate parameters in
ANUCLIM combine data in an interaction-like relationship (e.g., mean temperature of the wettest quarter) and clearly new variables that represent interactions
can be created for any of the methods, but there is no inbuilt capability for such
analysis.
There are a variety of approaches to working with correlated (collinear) predictor variables. ANUCLIM recommends the use of all climate parameters even
though the majority of them are often highly correlated. DOMAIN and GARP
have no recommended approach. Most texts on regression advise that dependencies among the covariates should be investigated (see, e.g., Hosmer and
Lemeshow 1989; Agresti 1996), but many statistical packages do not have inbuilt
routines for correlation analysis that are part of the modeling component. This
leaves it to the practitioners to develop their own good practice (see, e.g., Booth et
al. 1994).
Goodness of Fit
GLMs and GAMs are the only methods that provide statistical estimates of error
and deviance. These report the extent to which the model does not fit the data.
Usability and Transparency
All methods require some computing and statistical understanding. All can be
mastered with some effort, and none requires expert programming knowledge.
The time required to achieve proficiency will clearly vary with the user’s experience. The more important consideration is that the method and its output are well
understood by the user. Understanding is affected by the complexity of the underlying concepts, the transparency of the procedures, the form of the final model, the
extent of documentation of the particular programs used, the number of alternative packages for implementing the method (can the results be replicated?), and
the breadth of published applications of the method. Thus potential difficulties
with the more complex statistical concepts of the GLMs and GAMs are balanced
by the extensive documentation and literature associated with them (although
GAMs are much newer and are less commonly implemented), and the user-
