to aggregate individual stressor variables into a comprehensive index of impairment. Metrics and indices can then be evaluated as to the strength of their
correlation with this aggregate stressor index [124, 127, 128]. Tests of whether
metrics differ significantly from reference conditions provide binary results
regarding whether a metric classifies sites correctly. Standard approaches such
as ANOVA, however, may indicate statistically significant differences that are
not biologically meaningful [123, 129]; therefore, specialized techniques have
been developed that are more practical for determining whether metrics differ
from reference conditions [129–131]. The magnitude of the departure from
reference conditions is most important, regardless of statistical significance.
To this end, test statistics such as ANOVA F-statistics or t-scores, rather than
p-values, are used to determine the degree to which metrics differentiate
between reference and stressed sites [116, 123, 125]. Estimates of Type II
error rates are given by choosing a threshold value in the reference site
distribution that indicates impairment (e.g., 5th or 25th percentile for metrics
that decrease with stress and the 95th or 75th percentile for those that increase
with stress) and determining the proportion of impaired sites where metric
scores exceed this threshold (for metrics that increase with stress), indicating
that impairment has not been correctly identified [132]. Barbour et al. [133]
developed a similar, graphical approach for evaluating the degree to which
metrics and indices discriminate between reference and impaired conditions
(Fig. 4). Distribution-based methods such as these are especially susceptible to
the confounding effects of outliers, which should be carefully scrutinized to
determine whether they are caused by imprecise metrics or site misclassification.
Measures of precision describe the reliability of metrics and indices for consistently indicating site conditions. Those that exhibit high variability that is not
attributable to environmental predictors are not useful for bioassessment. Precision
is expressed by measures of variability in metric or index values among samples,
most commonly as the standard deviation (SD) or coefficient of variation (CV).
Variance partitioning is conducted to determine the relative importance of the three
primary sources of variation: among-site spatial variation, within-site spatial variation, and temporal variation [29, 134].
Temporal precision is often evaluated using the signal-to-noise ratio (S/N ) [135],
which is the ratio of metric variance among sites to variance among multiple visits
at the same site. When evaluated using both stressed and reference sites, S/N reflects
both accuracy and temporal precision. Stevenson et al. [125] set S/N > 2 as the
acceptable ratio for diatom metrics. Stoddard et al. [123] indicated that acceptable
S/N values should vary, based on organisms’ generation times, from >1 for algae to
>4 for fish (though these preliminary guidelines require further evaluation).
Within-site spatial precision is reflected by metric variability in samples collected
at the same site and time, which may be affected by sampling error among spatially
or temporally replicated samples [134, 136] or by variation among bioassessments
employing different protocols [134, 137].
250
A.L. Garey and L.A. Smock
Précédent

- 267/309

Suivant