4. Quantitative Methods for Modeling Species Habitat
53
In the present study, the data were not particularly well suited to the modeling
exercise. The 3,522 quadrats were compiled over many decades in response to a
variety of research agendas and management objectives. The error in spatial
location alone means that the finest scale at which they can confidently be used
(250 m) is too coarse for a good representation of site topographic conditions
(Elith et al. 1998). In my experience, the kinds of errors and gaps in the data used
in this exercise are more representative of the kind of information likely to be
available to management agencies than are the data collected specifically for
modeling studies.
Under these data conditions, there was no clear distinction in modeling success
between the methods. Variation in success between species was often greater than
variation in success between methods within species (Fig. 4.2). The errors in the
data were such that they masked most consistent differences between the performances of the methods. The lack of a clear difference between the methods is
probably as much a reflection on the general lack of discrimination of all models
as it is a revealing comparison of the methods. When data sets have more inherent
“noise,” there appears to be less advantage in using statistical models. Although
there may be advantages in using GAMs to model species habitat, a result that is
apparent in most comparative studies and even in these data, management agencies need to be aware that effort may be better invested in improving data quality
than in more sophisticated modeling techniques. However, there is still sometimes
a preference for the use of GLMs and GAMs because researchers like to be able to
interpret the models on an ecological level, and the relationship expressed as a
linear or additive model is easier to assess than, say, the complex rules evolved in
a genetic algorithm.
It is apparent that the models for some species have good enough discrimination to be useful at a subcatchment scale, but it is difficult to predict a priori which
species will be successfully modeled. This means that critical evaluation of any
model is essential to its proper application. Data sets for testing models are
commonly either the same as the modeling set or are partitioned subsets of the full
data set (see Fielding and Bell 1997 for a summary of data partitioning methods).
However, validation is more thorough if applied to a completely new data set
(Chatfield 1995). In this study, the conclusions about modeling success will be
dependent on the structure of the validation set. Issues of scale and scarcity need
to be carefully considered, so that the validation results reflect the types of
comparisons required by the models (Elith and Burgman, in review).
Variable selection is important, but not only in terms of the number of variables.
Also important is the set of variables. In this study, the set was constrained by
what was available. All methods would have done better with a set of variables
that reflected the habitat requirements of the particular species being modeled.
The competing argument is that in many circumstances, there will be no information on which to judge a priori which variables will be important. The problem of
variable selection is frequently discussed in relation to the statistical methods
(GLMs and GAMs), and specific approaches are advised for reducing the variable
set (see, e.g., Harrell and Lee 1984; Booth et al. 1994; Ferrier and Watson 1996;
53
In the present study, the data were not particularly well suited to the modeling
exercise. The 3,522 quadrats were compiled over many decades in response to a
variety of research agendas and management objectives. The error in spatial
location alone means that the finest scale at which they can confidently be used
(250 m) is too coarse for a good representation of site topographic conditions
(Elith et al. 1998). In my experience, the kinds of errors and gaps in the data used
in this exercise are more representative of the kind of information likely to be
available to management agencies than are the data collected specifically for
modeling studies.
Under these data conditions, there was no clear distinction in modeling success
between the methods. Variation in success between species was often greater than
variation in success between methods within species (Fig. 4.2). The errors in the
data were such that they masked most consistent differences between the performances of the methods. The lack of a clear difference between the methods is
probably as much a reflection on the general lack of discrimination of all models
as it is a revealing comparison of the methods. When data sets have more inherent
“noise,” there appears to be less advantage in using statistical models. Although
there may be advantages in using GAMs to model species habitat, a result that is
apparent in most comparative studies and even in these data, management agencies need to be aware that effort may be better invested in improving data quality
than in more sophisticated modeling techniques. However, there is still sometimes
a preference for the use of GLMs and GAMs because researchers like to be able to
interpret the models on an ecological level, and the relationship expressed as a
linear or additive model is easier to assess than, say, the complex rules evolved in
a genetic algorithm.
It is apparent that the models for some species have good enough discrimination to be useful at a subcatchment scale, but it is difficult to predict a priori which
species will be successfully modeled. This means that critical evaluation of any
model is essential to its proper application. Data sets for testing models are
commonly either the same as the modeling set or are partitioned subsets of the full
data set (see Fielding and Bell 1997 for a summary of data partitioning methods).
However, validation is more thorough if applied to a completely new data set
(Chatfield 1995). In this study, the conclusions about modeling success will be
dependent on the structure of the validation set. Issues of scale and scarcity need
to be carefully considered, so that the validation results reflect the types of
comparisons required by the models (Elith and Burgman, in review).
Variable selection is important, but not only in terms of the number of variables.
Also important is the set of variables. In this study, the set was constrained by
what was available. All methods would have done better with a set of variables
that reflected the habitat requirements of the particular species being modeled.
The competing argument is that in many circumstances, there will be no information on which to judge a priori which variables will be important. The problem of
variable selection is frequently discussed in relation to the statistical methods
(GLMs and GAMs), and specific approaches are advised for reducing the variable
set (see, e.g., Harrell and Lee 1984; Booth et al. 1994; Ferrier and Watson 1996;
