80
Sampling Design and Statistical Inference for Ecological Assessment
and almost always problematic to determine the
minimal sample size necessary for accurate estimates and sensitive tests. Sample design and the
statistical analysis of sample data may benefit from
the assistance or review of a trained statistician. In
recognition of the diversity and complexity of sampling problems in ecological assessment, the objective of this chapter is to review relevant statistical concepts and discuss the general principles of
sampling design so that researchers will be familiar with the issues and conversant in the language
of sampling design and statistical inference.
6.2 Statistical Inference
We always risk drawing an incorrect conclusion
when analyzing a popUlation based on a sample,
because not all population units are sampled. The
goal of statistical inference is to minimize the risk
of inferential errors and to provide a statement
about the quality of the inference. Drawing on the
previous example, a 95% confidence interval for
the proportion of infected stands in a watershed is
defined by its upper and lower bounds, U and L. A
good sampling design will produce intervals of
small width W, where W = U - L, and will ensure
that the intervals are actually correct 95% of the
time.
6.2.1 Populations and
Probability Samples
A population is a collection of elements, objects,
or organisms to which the findings of the study are
to be extrapolated. The elements of the population
are called units and the number of units in the population is denoted by N. Each unit provides a single measurement on a particular variable, attribute,
or characteristic, although multiple variables may
be measured on each unit. Multiple measurements
on a single unit are often thought of as a single observation on a multivariate variable. We denote the
variable of interest by Y, and a measurement of Y
made on a population unit is denoted by y. For example, if we have a map identifying forest stands
within a watershed, we may chose to define the
population units to be the stands and define Y to be
the triple: stand density, mean age of trees with
d.b.h. ;::: 12 inches, and infection rate of trees with
d.b.h. ;::: 12 inches. In principle, we can obtain a
single measurement on this multivariate variable
from each stand.
The elementary sampling setup assumes that N is
finite and known so that a list of labeled population
units can be formed. Other setups are also used, but
this one is useful for its simplicity. In this spirit, suppose that the y variable is univariate, and let YJ, Y2,
... , YN, denote the population of values. Parameters are descriptive characteristics of the popUlation
such as the population total, T = I1=1 Y;, the population mean, J. . L = TIN, and the population standard deviation, (T = VI~l (y; - J..L)2/N. The population total is a particularly useful parameter for
finite popUlations, because many other parameters
are simple functions of the population total, and
many estimators are functions of the sample total.
For example, an estimator of J. . L is the sample mean,
which is the sample total divided by the sample
size.
6.2.2 Probability Samples
A sample is a subset of the population that has been
selected by a researcher. A commonly used definition of a probability sample is one in which each
popUlation unit has a known, nonzero probability
of being included in the sample (Hansen et aI.,
1953; Levy and Lemeshow, 1991). The probability that a population unit is included in the sample
is called its inclusion probability. A statistically
valid sample is a probability sample; the two terms
are equivalent because probability sampling is
necessary for statistical inference. A list of the population units with their associated inclusion probabilities is called a sampling frame. Once a population unit has been included in the sample, it is called
a sample unit, and an observation is the value of
the y variable measured on a sample unit.
There are two elemental types of samplingwith and without replacement. A sample selected
without replacement is one in which every population unit can appear at most once in the sample,
whereas sampling with replacement allows each
unit to be selected more than once. Although the
idea of including more than one observation from
a population unit in the sample seems inefficient,
it simplifies some sampling protocols. For instance,
sampling with replication is appealing when estimating population size for a large animal by aerial
survey, because it may be difficult to determine
whether an animal has already been counted.
Simple Random Sampling
The simple random sample (SRS) design is the prototypical sampling design. In an SRS design that
employs sampling without replacement, every subset of the population of size n has the same probability of being selected as the sample. Rather than
listing and randomly selecting among all possible
subsets, it is more practical to collect an SRS by
Sampling Design and Statistical Inference for Ecological Assessment
and almost always problematic to determine the
minimal sample size necessary for accurate estimates and sensitive tests. Sample design and the
statistical analysis of sample data may benefit from
the assistance or review of a trained statistician. In
recognition of the diversity and complexity of sampling problems in ecological assessment, the objective of this chapter is to review relevant statistical concepts and discuss the general principles of
sampling design so that researchers will be familiar with the issues and conversant in the language
of sampling design and statistical inference.
6.2 Statistical Inference
We always risk drawing an incorrect conclusion
when analyzing a popUlation based on a sample,
because not all population units are sampled. The
goal of statistical inference is to minimize the risk
of inferential errors and to provide a statement
about the quality of the inference. Drawing on the
previous example, a 95% confidence interval for
the proportion of infected stands in a watershed is
defined by its upper and lower bounds, U and L. A
good sampling design will produce intervals of
small width W, where W = U - L, and will ensure
that the intervals are actually correct 95% of the
time.
6.2.1 Populations and
Probability Samples
A population is a collection of elements, objects,
or organisms to which the findings of the study are
to be extrapolated. The elements of the population
are called units and the number of units in the population is denoted by N. Each unit provides a single measurement on a particular variable, attribute,
or characteristic, although multiple variables may
be measured on each unit. Multiple measurements
on a single unit are often thought of as a single observation on a multivariate variable. We denote the
variable of interest by Y, and a measurement of Y
made on a population unit is denoted by y. For example, if we have a map identifying forest stands
within a watershed, we may chose to define the
population units to be the stands and define Y to be
the triple: stand density, mean age of trees with
d.b.h. ;::: 12 inches, and infection rate of trees with
d.b.h. ;::: 12 inches. In principle, we can obtain a
single measurement on this multivariate variable
from each stand.
The elementary sampling setup assumes that N is
finite and known so that a list of labeled population
units can be formed. Other setups are also used, but
this one is useful for its simplicity. In this spirit, suppose that the y variable is univariate, and let YJ, Y2,
... , YN, denote the population of values. Parameters are descriptive characteristics of the popUlation
such as the population total, T = I1=1 Y;, the population mean, J. . L = TIN, and the population standard deviation, (T = VI~l (y; - J..L)2/N. The population total is a particularly useful parameter for
finite popUlations, because many other parameters
are simple functions of the population total, and
many estimators are functions of the sample total.
For example, an estimator of J. . L is the sample mean,
which is the sample total divided by the sample
size.
6.2.2 Probability Samples
A sample is a subset of the population that has been
selected by a researcher. A commonly used definition of a probability sample is one in which each
popUlation unit has a known, nonzero probability
of being included in the sample (Hansen et aI.,
1953; Levy and Lemeshow, 1991). The probability that a population unit is included in the sample
is called its inclusion probability. A statistically
valid sample is a probability sample; the two terms
are equivalent because probability sampling is
necessary for statistical inference. A list of the population units with their associated inclusion probabilities is called a sampling frame. Once a population unit has been included in the sample, it is called
a sample unit, and an observation is the value of
the y variable measured on a sample unit.
There are two elemental types of samplingwith and without replacement. A sample selected
without replacement is one in which every population unit can appear at most once in the sample,
whereas sampling with replacement allows each
unit to be selected more than once. Although the
idea of including more than one observation from
a population unit in the sample seems inefficient,
it simplifies some sampling protocols. For instance,
sampling with replication is appealing when estimating population size for a large animal by aerial
survey, because it may be difficult to determine
whether an animal has already been counted.
Simple Random Sampling
The simple random sample (SRS) design is the prototypical sampling design. In an SRS design that
employs sampling without replacement, every subset of the population of size n has the same probability of being selected as the sample. Rather than
listing and randomly selecting among all possible
subsets, it is more practical to collect an SRS by
