weighted estimate from a subsample of observations in which the data are
present are used; adjust the definition of the target population; impute
the missing or removed data points; and explicitly model the missing data
(Little and Rubin 1983). The third and fourth methods approach the
problem from a sampling-theoretic paradigm, while the fifth and sixth
approach the problem from a modeling perspective. A statistician familiar
with the data-gathering method should be consulted for recommendations
and analysis if missing values are an issue. The last two methods attempt to
model the missing data, and one or more assumptions must be made as to
the underlying probability distribution of the data. Simple imputation has
subtle problems, which are not apparent at first glance. For instance, in the
simplest case of imputation, a mean value may be estimated. Although this
process may not appear to affect the sample mean, it may still affect the
estimate for the population and may bias the sample variance (Chernick
1983). Little and Rubin (1983) provide several situations under which data
may be missing and suggest imputation techniques to deal with them. In
some cases, multiple imputation methods (Little and Rubin 1987) provide
better ways of modeling the data without directly imputing individual
missing data points. Note that some statistical analysis techniques are relatively robust to missing data. For example, the mixed models methodology,
such as is incorporated in SAS PROC MIXED, is robust to random missing
observations in a multivariate, repeated measures context (Littell et al.
1996).
10.4.4 Autocorrelation
Autocorrelation occurs when samples (in either time or space) are more
like neighboring samples than distant samples. In such a case, the unit
of measurement (either time or distance) between sampling events is an
important predictor variable. In the temporal sense, this variation in similarity may be caused by an autoregressive process, where the data are correlated but the correlation decreases as the time between observations
increases. Autoregressive processes in environmental data can be even
subtler because they may have seasonal patterns embedded within them.
Time plots can sometimes help to show such features. Assuming that data
are normally independently and identically distributed, when in fact they
are not, can cause serious errors in hypothesis testing and extrapolation.
The existence of spatial dependence acts to reduce the degrees of freedom,
in effect decreasing the number of independent observations (Cressie
1993). This loss in degrees of freedom will increase the confidence intervals
of a prediction and reduce the ability to extrapolate model results. Metadata may reveal whether autocorrelation is a problem. For example, the
pilot data of the Forest Health Monitoring Program showed researchers
how far apart to place subplots and measurement points within subplots for
10. Effective Ecological Modeling: Data Issues
197
Précédent

- 203/327

Suivant