3.2 Advances in Predictive Modeling
The widespread application of RIVPACS-type models has inspired many alternative approaches. Recognizing that assemblages occur along continuous environmental gradients, investigators have developed nearest neighbor methods that
compare the environmental similarity of test sites to each individual reference
site, rather than to the average assemblage of each class as is done using RIVPACS
[83, 95]. Modeling approaches often skip the biotic classification step and predict
assemblage characteristics at reference sites directly using natural environmental
variables [96–98]. Direct prediction approaches may allow for different sets of
environmental variables to be used as predictors for each taxon. Though appealing
in this respect, the development of separate models for each taxon may be overly
complex for taxon-rich systems.
In contrast to the long history of predictive modeling for O/E indices [70], until
recently, developers of MMIs rarely employed predictive modeling to account for
natural environmental variability. McCormick et al. [99] used linear regression to
control for the effects of watershed size on a fish MMI. Equations derived from the
regression of metric values on watershed size at reference sites were applied to test
sites, and the residuals from the regression were used to indicate deviations from the
expected metric values in the absence of impairment. Oberdorff et al. [100]
expanded this approach, modeling metrics based on a suite of natural environmental
variables using logistic regression (for presence/absence metrics) and multiple
linear regression (for abundance-based metrics). Variations on this residualization
technique have been developed for more advanced modeling strategies such as
prediction tree approaches (discussed below), improving both the accuracy and
precision of MMIs by removing the confounding effects of natural environmental
variables [21, 101, 102].
Although conventional techniques such as MDA and linear and logistic regression have provided utility for predictive modeling, several newer methods better
account for the variable, often nonlinear and interactive effects of environmental
predictors on biota. The generalized additive modeling approach of Yuan [103]
shows the flexibility of this nonparametric regression technique for predicting
variable responses among different taxa to a suite of environmental factors. Bayesian frameworks provide a comprehensive evaluation of uncertainty in predictive
models [104–106] and have been used for this purpose in MMI development.
Machine learning techniques, including artificial neural networks (ANNs) and
ensemble prediction trees, where models are iteratively trained at prediction to
minimize error, have received much recent attention for predictive modeling in
ecology [107–110]. Though the method is not yet widely used, support vector
machines have performed favorably compared with other machine learning techniques for predicting the occurrence of macroinvertebrate taxa [111, 112].
ANNs structure predictor–response relationships in a manner similar to vertebrate neurological systems. Variables are represented as neurons connected by a
multitude of axons representing the possible interrelationships among variables
Principles for the Development of Contemporary Bioassessment Indices for. . .
247
The widespread application of RIVPACS-type models has inspired many alternative approaches. Recognizing that assemblages occur along continuous environmental gradients, investigators have developed nearest neighbor methods that
compare the environmental similarity of test sites to each individual reference
site, rather than to the average assemblage of each class as is done using RIVPACS
[83, 95]. Modeling approaches often skip the biotic classification step and predict
assemblage characteristics at reference sites directly using natural environmental
variables [96–98]. Direct prediction approaches may allow for different sets of
environmental variables to be used as predictors for each taxon. Though appealing
in this respect, the development of separate models for each taxon may be overly
complex for taxon-rich systems.
In contrast to the long history of predictive modeling for O/E indices [70], until
recently, developers of MMIs rarely employed predictive modeling to account for
natural environmental variability. McCormick et al. [99] used linear regression to
control for the effects of watershed size on a fish MMI. Equations derived from the
regression of metric values on watershed size at reference sites were applied to test
sites, and the residuals from the regression were used to indicate deviations from the
expected metric values in the absence of impairment. Oberdorff et al. [100]
expanded this approach, modeling metrics based on a suite of natural environmental
variables using logistic regression (for presence/absence metrics) and multiple
linear regression (for abundance-based metrics). Variations on this residualization
technique have been developed for more advanced modeling strategies such as
prediction tree approaches (discussed below), improving both the accuracy and
precision of MMIs by removing the confounding effects of natural environmental
variables [21, 101, 102].
Although conventional techniques such as MDA and linear and logistic regression have provided utility for predictive modeling, several newer methods better
account for the variable, often nonlinear and interactive effects of environmental
predictors on biota. The generalized additive modeling approach of Yuan [103]
shows the flexibility of this nonparametric regression technique for predicting
variable responses among different taxa to a suite of environmental factors. Bayesian frameworks provide a comprehensive evaluation of uncertainty in predictive
models [104–106] and have been used for this purpose in MMI development.
Machine learning techniques, including artificial neural networks (ANNs) and
ensemble prediction trees, where models are iteratively trained at prediction to
minimize error, have received much recent attention for predictive modeling in
ecology [107–110]. Though the method is not yet widely used, support vector
machines have performed favorably compared with other machine learning techniques for predicting the occurrence of macroinvertebrate taxa [111, 112].
ANNs structure predictor–response relationships in a manner similar to vertebrate neurological systems. Variables are represented as neurons connected by a
multitude of axons representing the possible interrelationships among variables
Principles for the Development of Contemporary Bioassessment Indices for. . .
247
