reduce the mean correlation among a group of potential metrics, multivariate
analyses such as PCA and cluster analysis can be used to aggregate correlated
metrics [21, 101, 102]. The correlation of metric errors (e.g., residuals of stressor–
metric regressions) rather than the correlation of metric values may be more
appropriate for judging redundancy, a concept whose applicability should be
further evaluated [140].
4.4 Metric Aggregation and Scoring
MMI development is completed by aggregating the best-performing metrics to
derive an index score. The number of metrics used in the index varies among
studies, and the choice is rarely supported by clear empirical justification [139].
Professional judgment is often used to select metrics based on best overall performance, although ordered stepwise processes have been recommended and present
more comprehensive and objective options [139, 140]. To express metrics on an
equivalent numerical scale, raw values are commonly rescaled to reflect percent or
proportional comparability to values in the reference site distribution or to the
distribution of all sites producing metric scores on continuous 0–100 or 0–1 point
scales that increase with impairment. Blocksom [49] reviewed the details of these
and other common scoring methods. After scoring, metrics are nearly always
aggregated into an index by simple averaging, although other methods, such as
differential weighting based on relative importance [131, 141] or to account for
variations in metric precision [131], have been used. Alternative aggregation
strategies for MMIs represent yet another area where additional research is needed.
4.5 Index Validation
Validation of the index with independent data provides the most comprehensive
evaluation of performance. Validation typically proceeds by randomly selecting
subsets of impaired and reference sites, which are excluded from the dataset used
for index development and used for a posteriori evaluation of the performance
characteristics described above. The feasibility of index validation depends on the
amount of data available, as statistical power is compromised by dividing datasets
for this purpose. Categorical approaches for validating index accuracy are data
expensive, as the validation set must be divided according to impairment status.
When only a few sites are available, index accuracy may be validated by analyzing
for correlations of index scores with stressor gradients, which requires fewer
validation sites. This approach is especially useful in highly developed landscapes
where there are few reference sites [127, 142]. Index accuracy is often prioritized
over other performance characteristics, although more thorough validation strategies also evaluate precision [21, 106, 143]. Evaluation of index bias, as indicated by
Principles for the Development of Contemporary Bioassessment Indices for. . .
253
analyses such as PCA and cluster analysis can be used to aggregate correlated
metrics [21, 101, 102]. The correlation of metric errors (e.g., residuals of stressor–
metric regressions) rather than the correlation of metric values may be more
appropriate for judging redundancy, a concept whose applicability should be
further evaluated [140].
4.4 Metric Aggregation and Scoring
MMI development is completed by aggregating the best-performing metrics to
derive an index score. The number of metrics used in the index varies among
studies, and the choice is rarely supported by clear empirical justification [139].
Professional judgment is often used to select metrics based on best overall performance, although ordered stepwise processes have been recommended and present
more comprehensive and objective options [139, 140]. To express metrics on an
equivalent numerical scale, raw values are commonly rescaled to reflect percent or
proportional comparability to values in the reference site distribution or to the
distribution of all sites producing metric scores on continuous 0–100 or 0–1 point
scales that increase with impairment. Blocksom [49] reviewed the details of these
and other common scoring methods. After scoring, metrics are nearly always
aggregated into an index by simple averaging, although other methods, such as
differential weighting based on relative importance [131, 141] or to account for
variations in metric precision [131], have been used. Alternative aggregation
strategies for MMIs represent yet another area where additional research is needed.
4.5 Index Validation
Validation of the index with independent data provides the most comprehensive
evaluation of performance. Validation typically proceeds by randomly selecting
subsets of impaired and reference sites, which are excluded from the dataset used
for index development and used for a posteriori evaluation of the performance
characteristics described above. The feasibility of index validation depends on the
amount of data available, as statistical power is compromised by dividing datasets
for this purpose. Categorical approaches for validating index accuracy are data
expensive, as the validation set must be divided according to impairment status.
When only a few sites are available, index accuracy may be validated by analyzing
for correlations of index scores with stressor gradients, which requires fewer
validation sites. This approach is especially useful in highly developed landscapes
where there are few reference sites [127, 142]. Index accuracy is often prioritized
over other performance characteristics, although more thorough validation strategies also evaluate precision [21, 106, 143]. Evaluation of index bias, as indicated by
Principles for the Development of Contemporary Bioassessment Indices for. . .
253
