4. Quantitative Methods for Modeling Species Habitat
49
to try to improve the probability of finding the rare species, and the remaining
sites were selected randomly from locations not on private land, within the constraints of the criteria: 391 sites were selected. At each, the following species
information was recorded:
• presence of each of the species within a 30 × 30-m quadrat located centrally
within a 250 × 250-m target cell
• presence of each of the species anywhere within the cell
• presence of habitat suitable for the species (as assessed by the botanist) within
the 30 × 30-m quadrat (“habitat probable”)
• presence of habitat that is possible but not optimal for the species within the
quadrat, or habitat suitable for the species outside the quadrat but still within the
cell (“habitat possible”).
Statistical Analysis
In this case, in which different forms of predictions (probabilities, relative likelihoods, and ranks) are produced and in which the ability of the models to predict
distribution in new unsampled areas is critical, a measure of discrimination will
provide the best comparison of methods. Confusion matrices are commonly used
to assess this type of accuracy, but they require the selection of a threshold value
that defines whether a prediction is for presence or absence. This is not necessarily
a straightforward task (see Fielding and Bell 1997). Receiver operating characteristic (ROC) curves are threshold independent, because they assess the truepositive fraction against the false-positive fraction over numerous decision thresholds (Metz 1978). The area under the ROC curve indicates the percentage of
correct decisions in paired comparisons (Swets 1988), which in a species modeling context will be an estimate of the probability of correctly ranking a presenceabsence pair. For example, an area of 0.78 would indicate that, given a randomly
selected pair of presence and absence observations, the model prediction at the
presence site would be higher than at the absence site 78% of the time. This ROC
area is equivalent to Wilcoxon’s Mann-Whitney version of the nonparameteric
two-sample statistic (Hanley and McNeil 1982; Ferrier and Watson 1996). The
ROC area is more complex to compute and requires specialist software (for
software listings, see Zweig and Campbell 1992; S-News 1999), but the MannWhitney statistic is a satisfactory alternative and is easy to compute. Standard
errors can be calculated for either statistic (Zweig and Campbell 1992; Hanley and
McNeil 1982; DeLong et al. 1988); the method of DeLong and associates (1988)
was used in this study. The ROC area can vary from 0 to 1, although the convention in the medical field in which it originated is to constrain it to 0.5 or greater by
reversing the decision rule if it is less than 0.5. This is not appropriate in the
current context. An ROC area of 1.0 indicates perfect discrimination, whereas 0.5
indicates that discrimination is equivalent to that of a random set of predictions.
Here, we view models with an ROC area greater than 0.75 as ones with sufficient
discrimination to be potentially useful in reserve planning.
49
to try to improve the probability of finding the rare species, and the remaining
sites were selected randomly from locations not on private land, within the constraints of the criteria: 391 sites were selected. At each, the following species
information was recorded:
• presence of each of the species within a 30 × 30-m quadrat located centrally
within a 250 × 250-m target cell
• presence of each of the species anywhere within the cell
• presence of habitat suitable for the species (as assessed by the botanist) within
the 30 × 30-m quadrat (“habitat probable”)
• presence of habitat that is possible but not optimal for the species within the
quadrat, or habitat suitable for the species outside the quadrat but still within the
cell (“habitat possible”).
Statistical Analysis
In this case, in which different forms of predictions (probabilities, relative likelihoods, and ranks) are produced and in which the ability of the models to predict
distribution in new unsampled areas is critical, a measure of discrimination will
provide the best comparison of methods. Confusion matrices are commonly used
to assess this type of accuracy, but they require the selection of a threshold value
that defines whether a prediction is for presence or absence. This is not necessarily
a straightforward task (see Fielding and Bell 1997). Receiver operating characteristic (ROC) curves are threshold independent, because they assess the truepositive fraction against the false-positive fraction over numerous decision thresholds (Metz 1978). The area under the ROC curve indicates the percentage of
correct decisions in paired comparisons (Swets 1988), which in a species modeling context will be an estimate of the probability of correctly ranking a presenceabsence pair. For example, an area of 0.78 would indicate that, given a randomly
selected pair of presence and absence observations, the model prediction at the
presence site would be higher than at the absence site 78% of the time. This ROC
area is equivalent to Wilcoxon’s Mann-Whitney version of the nonparameteric
two-sample statistic (Hanley and McNeil 1982; Ferrier and Watson 1996). The
ROC area is more complex to compute and requires specialist software (for
software listings, see Zweig and Campbell 1992; S-News 1999), but the MannWhitney statistic is a satisfactory alternative and is easy to compute. Standard
errors can be calculated for either statistic (Zweig and Campbell 1992; Hanley and
McNeil 1982; DeLong et al. 1988); the method of DeLong and associates (1988)
was used in this study. The ROC area can vary from 0 to 1, although the convention in the medical field in which it originated is to constrain it to 0.5 or greater by
reversing the decision rule if it is less than 0.5. This is not appropriate in the
current context. An ROC area of 1.0 indicates perfect discrimination, whereas 0.5
indicates that discrimination is equivalent to that of a random set of predictions.
Here, we view models with an ROC area greater than 0.75 as ones with sufficient
discrimination to be potentially useful in reserve planning.
