48
Jane Elith
ANUCLIM used climate predictor variables and species presence data. Models
developed in DOMAIN and GARP were constructed with the full set of environmental variables. Because there is noticable variation between runs of the genetic
algorithm, spatial predictions were the average of ten iterations of the program.
GLMs were fitted to the data as logistic regression models on a binary response
with a logit link function (Hosmer and Lemeshow 1989). All available predictor
variables were initial candidates, with vegetation type handled as separate binary
variables. For each species, the predictor candidate set was reduced to a subset
through a univariate logistic analysis, in which a variable was retained for the
subset if the change in deviance between the model containing the variable and the
null model was significant at a p value of .2. In the subset, variables that were
correlated were excluded through analysis of variance inflation factors (Booth et
al. 1994; Sokal and Rohlf 1981). Models were developed with an automated
forward selection procedure in which, at each step, the variable associated with
the largest reduction in deviance was retained, and the process was repeated until
none of the remaining variables significantly improved the model (p < .05). If
latitude and longitude were included in the subset, they were considered last.
Linear and quadratic terms were investigated for all continuous variables. Interactions were not considered. Most of these protocols were developed for species
modeling being used by the government in land-use planning, and part of the
intention of the study was to investigate modeling success under these protocols.
Variable selection for GAMs was approached in the same manner as for GLMs.
GAMs were developed with 4 degrees of freedom initially assigned to each
continuous variable, with subsequent testing of changes in deviance used to
reduce the degrees of freedom when possible. Relationships were smoothed with
cubic splines. Models for GLMs and GAMs were developed using S-PLUS (version 3.3 for Windows 95, MathSoft 1995).
Validation
Validation of the predictions was based on field surveys carried out by an experienced botanist in 1996 and 1997. Consistency between the original survey data
and the validation set was achieved by using the same basic survey unit, a 900-m
2
quadrat, which was usually 30 × 30 m, or 15 × 60 m in riparian habitat. The
botanist recorded the presence or absence of all 29 species, and searching continued until there was a reasonable certainty that any of the species would have
been sighted if they were present. Quadrats were located on 1:100,000 mapsheets,
and when these mapsheets did not appear consistent with the observed topography, 1:25,000 sheets were used to locate quadrats more precisely.
Validation sites were selected on the basis of model predictions. Two criteria
were addressed: at least three sites from the top 0.1% of the distribution of
predicted values for each species and each method were required, and for each
species/method combination the full set of sites should span the full range of
predictions. The choice of the first 100 sites was conditioned by expert judgment,
Jane Elith
ANUCLIM used climate predictor variables and species presence data. Models
developed in DOMAIN and GARP were constructed with the full set of environmental variables. Because there is noticable variation between runs of the genetic
algorithm, spatial predictions were the average of ten iterations of the program.
GLMs were fitted to the data as logistic regression models on a binary response
with a logit link function (Hosmer and Lemeshow 1989). All available predictor
variables were initial candidates, with vegetation type handled as separate binary
variables. For each species, the predictor candidate set was reduced to a subset
through a univariate logistic analysis, in which a variable was retained for the
subset if the change in deviance between the model containing the variable and the
null model was significant at a p value of .2. In the subset, variables that were
correlated were excluded through analysis of variance inflation factors (Booth et
al. 1994; Sokal and Rohlf 1981). Models were developed with an automated
forward selection procedure in which, at each step, the variable associated with
the largest reduction in deviance was retained, and the process was repeated until
none of the remaining variables significantly improved the model (p < .05). If
latitude and longitude were included in the subset, they were considered last.
Linear and quadratic terms were investigated for all continuous variables. Interactions were not considered. Most of these protocols were developed for species
modeling being used by the government in land-use planning, and part of the
intention of the study was to investigate modeling success under these protocols.
Variable selection for GAMs was approached in the same manner as for GLMs.
GAMs were developed with 4 degrees of freedom initially assigned to each
continuous variable, with subsequent testing of changes in deviance used to
reduce the degrees of freedom when possible. Relationships were smoothed with
cubic splines. Models for GLMs and GAMs were developed using S-PLUS (version 3.3 for Windows 95, MathSoft 1995).
Validation
Validation of the predictions was based on field surveys carried out by an experienced botanist in 1996 and 1997. Consistency between the original survey data
and the validation set was achieved by using the same basic survey unit, a 900-m
2
quadrat, which was usually 30 × 30 m, or 15 × 60 m in riparian habitat. The
botanist recorded the presence or absence of all 29 species, and searching continued until there was a reasonable certainty that any of the species would have
been sighted if they were present. Quadrats were located on 1:100,000 mapsheets,
and when these mapsheets did not appear consistent with the observed topography, 1:25,000 sheets were used to locate quadrats more precisely.
Validation sites were selected on the basis of model predictions. Two criteria
were addressed: at least three sites from the top 0.1% of the distribution of
predicted values for each species and each method were required, and for each
species/method combination the full set of sites should span the full range of
predictions. The choice of the first 100 sites was conditioned by expert judgment,
