temporal variability of sample measurements at a field site, and spatial
variability of quantities to be measured in the field. Sampling error can
sometimes be estimated, most commonly by the use of properly designed
quality-control programs with duplicate samples and/or replicate measurements. For example, the Forest Health Monitoring Program has successfully
used quality-control data in pilot and demonstration projects to estimate
the comparative magnitude of measurement error and hence to assess
whether specific measurements were reliable enough to use in further
program fieldwork. (Riiters et al. 1991.)
Nonsampling errors are all errors that are not sampling errors, generally
attributable to the manner in which observations are made and encompassing many of the practical problems of implementing a sample design.
Sources of nonsampling error include nonobservation, exclusion of certain
groups or subgroups, inclusion of inappropriate sites, difficulties with definitions, and differing interpretations of class information. Quantifying nonsampling errors can be difficult and expensive because it requires multiple
observations of the same phenomena. For this reason, nonsampling errors
often go unnoticed. The most reasonable approach to nonsampling errors
is to employ a good statistician to design the experiment or sample frame
and to use an experienced researcher as an auditor. The Environmental
Monitoring and Assessment Program (EMAP) and the Forest Health Monitoring Program are good examples of the careful use of a well-designed
sample-site selection method based on rigorous probability-sampling techniques, and both have demonstrated the ability to adjust sample weights to
handle nonsampling errors (Overton et al. 1990; Palmer et al. 1992). When
using other people’s data, it is imperative to have good metadata to assist
in determining whether the source of the data is considered to be reliable.
Quality-control data can provide information on the reliability of the data
set, but are often difficult to acquire, even when they exist. Statisticians may
sometimes be able to use such metadata to evaluate the data or (in rare
cases) assess the validity of data points. As one example, the USEPA’s
Direct Delayed Response Project (DDRP) had a quality-control component sufficient to allow evaluation of individual data points so that each
point in its soil chemistry database has a quality assurance (QA) flag indicating the quality of that data point (Van Remortel et al. 1988).
If no QA data or metadata are available, it may be impossible to determine the validity of any point, even apparent outliers. Still, in some cases,
the data may be determined from the metadata to be from a single welldefined population. In such a case, standard statistical techniques may be
of use for checking the data and assessing outliers. However, such is generally not the case with synoptic data or data collected over long time spans.
Plots of the raw data can also be extremely useful. However, care must be
taken because the data may represent a mixture of distributions, so that
simple assessments based on normality (or other common distributions)
may lead to problems.
10. Effective Ecological Modeling: Data Issues
195
Précédent

- 201/327

Suivant