4 Index Development and Performance Evaluation
We begin this section by discussing methods for evaluating the performance of
metrics and indices (Sects. 4.1–4.4). For MMIs, these characteristics should be
evaluated in order to include the best-performing metrics in the final index. Final
metric selection, scoring, and aggregation are discussed in Sect. 4.5. After scoring
and aggregation of metrics within an MMI or alternatively the development of an
O/E index, performance should be re-evaluated using the finished index scores,
ideally using independent data not used for index construction (Sect. 4.5). Further
information on MMI development has been presented by others [1, 32, 122]. For
clarity, these works present index development in a stepwise manner; however, it is
important to note that index development is an iterative, rather than a linear process.
Metrics that are acceptable based on one criterion (e.g., numerical range, Sect. 4.1)
may subsequently be considered unacceptable based on another criterion
(e.g., accuracy, Sect. 4.2), requiring the evaluation of new metrics.
4.1 Numerical Range
Assemblage data are often plagued with abundant zeros due to the patchy distribution of biota among habitats, and metrics related to rare taxa typically have narrow
numerical ranges. Metrics with limited ranges, and those for which many sites in
the dataset exhibit the same value, are unlikely to exhibit clear numerical responses
to stressors [123]. Others have presented guidelines for acceptable numerical ranges
for metrics, though these vary among studies [123–125]. Simple distribution plots
of metric values often provide clear indications of highly limited metrics (e.g., see
Fig. 2 in [122]).
4.2 Accuracy and Precision
We broadly define accuracy as the degree to which a given metric or index is
quantitatively related to variations in anthropogenic stress. Accurate metrics and
indices exhibit low Type II error rates by correctly identifying impairment and low
Type I error rates by correctly identifying reference conditions. As others have
indicated [126], the precise impairment state of a system, and therefore the absolute
accuracy of metrics, can never be truly known. We therefore use the term accuracy
to refer to estimated accuracy for identifying impairment, as indicated by relationships of metrics with a priori-selected stressor variables.
Relationships between metrics with continuously varying stressor variables
may be expressed using correlation analysis. The objective is often to assess the
responsiveness of metrics to overall stressor gradients. PCA is commonly used
Principles for the Development of Contemporary Bioassessment Indices for. . .
249
We begin this section by discussing methods for evaluating the performance of
metrics and indices (Sects. 4.1–4.4). For MMIs, these characteristics should be
evaluated in order to include the best-performing metrics in the final index. Final
metric selection, scoring, and aggregation are discussed in Sect. 4.5. After scoring
and aggregation of metrics within an MMI or alternatively the development of an
O/E index, performance should be re-evaluated using the finished index scores,
ideally using independent data not used for index construction (Sect. 4.5). Further
information on MMI development has been presented by others [1, 32, 122]. For
clarity, these works present index development in a stepwise manner; however, it is
important to note that index development is an iterative, rather than a linear process.
Metrics that are acceptable based on one criterion (e.g., numerical range, Sect. 4.1)
may subsequently be considered unacceptable based on another criterion
(e.g., accuracy, Sect. 4.2), requiring the evaluation of new metrics.
4.1 Numerical Range
Assemblage data are often plagued with abundant zeros due to the patchy distribution of biota among habitats, and metrics related to rare taxa typically have narrow
numerical ranges. Metrics with limited ranges, and those for which many sites in
the dataset exhibit the same value, are unlikely to exhibit clear numerical responses
to stressors [123]. Others have presented guidelines for acceptable numerical ranges
for metrics, though these vary among studies [123–125]. Simple distribution plots
of metric values often provide clear indications of highly limited metrics (e.g., see
Fig. 2 in [122]).
4.2 Accuracy and Precision
We broadly define accuracy as the degree to which a given metric or index is
quantitatively related to variations in anthropogenic stress. Accurate metrics and
indices exhibit low Type II error rates by correctly identifying impairment and low
Type I error rates by correctly identifying reference conditions. As others have
indicated [126], the precise impairment state of a system, and therefore the absolute
accuracy of metrics, can never be truly known. We therefore use the term accuracy
to refer to estimated accuracy for identifying impairment, as indicated by relationships of metrics with a priori-selected stressor variables.
Relationships between metrics with continuously varying stressor variables
may be expressed using correlation analysis. The objective is often to assess the
responsiveness of metrics to overall stressor gradients. PCA is commonly used
Principles for the Development of Contemporary Bioassessment Indices for. . .
249
