82
Sampling Design and Statistical Inference for Ecological Assessment
population as a fixed and usually finite set. Commonly, the objectives of sampling are to describe
the existing population units in terms of parameter
estimates and confidence intervals. The second
view is model based and is sometimes called a superpopulation view. A model-based view treats the
population as a single realization of a stochastic
process. In ecological terms, this is equivalent to
viewing the ecosystem as dynamic rather than static.
A model-based view is appropriate when the interest is in drawing inference not only about the
population units in existence at the time of sampling, but also about population units that existed
in the past or will exist in the future. A designbased view is sensible when sampling a forest stand
with the intent of estimating basal area or volume,
because the researcher is most likely interested in
what exists at the time of sampling. A model-based
view is more appropriate when the objective is to
monitor a riverine ecosystem by sampling for
changes in water chemistry or biotic components.
In this situation, a model describing the distribution of chemical constituents is useful for distinguishing between natural variation and trends in average concentration.
The statistical analysis of model-based samples
may differ from that of design-based samples; this
difference affects sampling design. For instance, a
routine objective of model-based sampling is to estimate a regression model describing a response
variable as a function of a set of explanatory variables. A model-based design may be aimed at selecting a sample with observations regularly distributed across the range of the explanatory
variable, because this design will tend to produce
more accurate estimates and model predictions than
a design that randomly samples the range. A modelbased design may stratify the population based on
the range of the explanatory variables and select an
independent SRS from each stratum. This design
will cover the range more fully than an SRS. A consequence of the model-based design is that some
estimates may be biased when examined from the
design view. For example, the sample mean of the
response variable may be biased under the modelbased design because subsamples have not been
proportionally allocated to strata. Overton and
Stehman (1995) provide further discussion relevant
to natural resource sampling, and Thompson (1992)
makes clear distinctions between the statistical
methods appropriate for the two views.
Statistical methods for analyzing data obtained
from both design- and model-based sampling can
be differentiated from a much larger set of statistical methods in two principal aspects. First, the statistical analysis of sampling data often assumes a
finite population. Second, and more importantly,
many designs sample observations that are not independent. Because statistical analysis of designed
samples is a large and difficult area, this chapter
discusses only elementary estimators in the context
of specific designs. The reader is referred to sampling design texts by Levy and Lemeshow (1991),
Siirndal et al. (1992), Schreuder et al. (1993), and
Thompson (1992) for statistical analysis methods for
data collected according to sampling designs. There
are, however, general principles guiding the use of
estimators, predictors, and hypothesis tests. The remainder of this section concentrates on these general principles, rather than specific methods. The finite-population, design-based view is adopted in the
remainder of this chapter, because this view clearly
departs from the conventional treatment of data as
independent observations from infinite populations.
6.4.1 Bias, Precision, and Accuracy
Further discussion of sampling design requires a
review of the properties of statistics, because sampling design affects the statistics used for inference
and hence the quality of the inference. Three of
these properties are bias, precision, and accuracy.
Suppose that a statistic S is used to estimate a parameter f The bias of S is the difference between
its expected value and g; that is, the bias of S is
E(S) - g. If the expected value of S is g, then S is
unbiased. An estimator need not be unbiased to be
useful; for instance, the sample standard deviation
is a biased estimator of the population standard deviation, but for practical applications the bias is
small enough to be ignored.
The precision of an estimator is a measure of the
repeatability of the estimator across many samples.
Reliability and reproducibility are equivalent to
precision. A precise estimator is reliable and reproducible in the sense that it produces similar estimates from different samples. Formally, the precision of an estimator is its variance, which is
defined to be the expected value of the mean
squared difference between the estimator and the
expected value of the estimator; that is, the variance of S is (T~ = E{[S - E(S)]2}. Suppose that m
possible values may be obtained from an estimator
S when all possible samples are considered. If the
set of possible values, or estimates, is denoted by
Sl, . . . , Sm, and Pi denotes the probability that S
will take on the value Si, then a computational formula for the variance is (T~ = If'!,l Pi [Si - E(s)]2.
The standard deviation (Ts of S, or the standard error of S, is (Ts = \Ici1.
Précédent

- 92/539

Suivant