6.6 Basic Sampling Designs
can be derived from T; for example, TIN is an estimator of the population mean.
The general form of the Horvitz-Thompson estimator is simple to define. Let 7T; denote the probability that the i th population unit will be included
in the sample. Then the Horvitz-Thompson estimator of the population total T is
T = f ll,
;=1 7T;
(1)
where Yb ... , Yv are the distinct observations in
the sample. If sampling occurs with replacement
and an observation appears more than once, only
the first appearance is used in the sum and counted
toward v. To illustrate the Horvitz-Thompson estimator, simple random sampling without replacement yields an inclusion probability of 7T; = nlN,
and the Horvitz-Thompson estimator becomes
T = (Nln) !,7=1 Y;, which is the usual sample mean
multiplied by the number of population units N. An
ecological example is that of estimating grass production. An SRS of plot locations can be selected
from the population of plots, and the grass on each
plot can be clipped, dried, and weighed. Mean production per square meter can be estimated by dividing T by the total number of square meters that
were sampled.
The power of the Horvitz-Thompson estimator
is demonstrated by the following example. Suppose
that the objective is to estimate the number of birds
in a fixed area by traversing a line transect and
counting the number of birds that are flushed. The
inclusion probabilities can be modeled as a function of the distance to the flushed birds (Hayne,
1949; Overton, 1969) according to the model 7T; =
2r;lL, where the ith bird will flush if the sampler is
within a distance of r; of the bird, and L is the transect length. Let y; be 1 for all i; then the sample
sum of the y;' s is the number of birds observed
while traversing the transect. The Horvitz-Thompson estimator of the population size is then N =
!,i'= 1 y;l7T; = (Ll2) !,i'= 1 lIr;, where v is the number of flushed birds.
There are two important points to be made about
this design. First, a bird is allowed to be observed
more than once, which is important because it is
often difficult to determine if the same bird has
been flushed more than once. Second, we do not
need to know the inclusion probabilities for all population units. This is important because we do not
even know the number of units, much less the inclusion probabilities. In fact, the often made statement that inclusion probabilities must be known for
all population units is stronger than necessary. In85
stead, the essential condition is that the inclusion
probabilities must be known for all units that appear in the sample; in other words, the inclusion
probabilities must be knowable for all population
units (Overton and Stehman, 1995) and nonzero.
The next three sections describe basic sampling designs used in ecological assessment.
6.6.1 Stratified Random Sampling
Stratified random sampling is a simple and effective method for improving on the efficiency of
simple random sampling. The idea is to divide the
population into a set of distinct strata, or subpopulations, so that the units within strata are more alike
than units between strata. If stratification has been
successful in this regard, then a smaller sample size
is necessary to achieve the same accuracy level than
if an SRS design were used. For example, a design
aimed at assessing stream condition can be made
more efficient if the population, or region of interest, is stratified according to watershed. Often, the
principal reason for stratification is to obtain strataspecific estimates, for example, when strata correspond to administrative units. Stratified random
sampling can be used to obtain both strata-specific
and population-wide estimates. An equally important motivation for stratified random sampling is
that the logistics of sampling are sometimes simplified, or the cost of sampling is reduced. For example, if a National Forest is to be sampled, it may
be efficient to stratify on a district basis, because
personnel located within each district can sample
that area at a lower cost than centrally located personnel. Cost and feasibility are adequate reasons
for stratified random sampling even if no gain in
accuracy is anticipated. However, it is possible to
degrade the quality of inferential methods if there
are many strata that are not particularly different,
and so it is best not to stratify unless there are
clearly different strata or clear advantages in terms
of feasibility or cost.
For stratified random sampling, the HorvitzThompson estimator is set up as follows. Suppose
that there are Nr popUlation units in the rth stratum. Stratified random sampling specifies that an
SRS of nr units be selected from the rth stratum.
Suppose that the ith population unit is a member
of the rth stratum. Then the inclusion probability
for the ith unit is 7T; = nrlNr. Equation (1) is the
Horvitz-Thompson estimator of the popUlation
total. Stratum-specific estimates are based on
estimates of the stratum total Tr for the rth strata.
The Horvitz-Thompson estimate of Tr is Tr =
can be derived from T; for example, TIN is an estimator of the population mean.
The general form of the Horvitz-Thompson estimator is simple to define. Let 7T; denote the probability that the i th population unit will be included
in the sample. Then the Horvitz-Thompson estimator of the population total T is
T = f ll,
;=1 7T;
(1)
where Yb ... , Yv are the distinct observations in
the sample. If sampling occurs with replacement
and an observation appears more than once, only
the first appearance is used in the sum and counted
toward v. To illustrate the Horvitz-Thompson estimator, simple random sampling without replacement yields an inclusion probability of 7T; = nlN,
and the Horvitz-Thompson estimator becomes
T = (Nln) !,7=1 Y;, which is the usual sample mean
multiplied by the number of population units N. An
ecological example is that of estimating grass production. An SRS of plot locations can be selected
from the population of plots, and the grass on each
plot can be clipped, dried, and weighed. Mean production per square meter can be estimated by dividing T by the total number of square meters that
were sampled.
The power of the Horvitz-Thompson estimator
is demonstrated by the following example. Suppose
that the objective is to estimate the number of birds
in a fixed area by traversing a line transect and
counting the number of birds that are flushed. The
inclusion probabilities can be modeled as a function of the distance to the flushed birds (Hayne,
1949; Overton, 1969) according to the model 7T; =
2r;lL, where the ith bird will flush if the sampler is
within a distance of r; of the bird, and L is the transect length. Let y; be 1 for all i; then the sample
sum of the y;' s is the number of birds observed
while traversing the transect. The Horvitz-Thompson estimator of the population size is then N =
!,i'= 1 y;l7T; = (Ll2) !,i'= 1 lIr;, where v is the number of flushed birds.
There are two important points to be made about
this design. First, a bird is allowed to be observed
more than once, which is important because it is
often difficult to determine if the same bird has
been flushed more than once. Second, we do not
need to know the inclusion probabilities for all population units. This is important because we do not
even know the number of units, much less the inclusion probabilities. In fact, the often made statement that inclusion probabilities must be known for
all population units is stronger than necessary. In85
stead, the essential condition is that the inclusion
probabilities must be known for all units that appear in the sample; in other words, the inclusion
probabilities must be knowable for all population
units (Overton and Stehman, 1995) and nonzero.
The next three sections describe basic sampling designs used in ecological assessment.
6.6.1 Stratified Random Sampling
Stratified random sampling is a simple and effective method for improving on the efficiency of
simple random sampling. The idea is to divide the
population into a set of distinct strata, or subpopulations, so that the units within strata are more alike
than units between strata. If stratification has been
successful in this regard, then a smaller sample size
is necessary to achieve the same accuracy level than
if an SRS design were used. For example, a design
aimed at assessing stream condition can be made
more efficient if the population, or region of interest, is stratified according to watershed. Often, the
principal reason for stratification is to obtain strataspecific estimates, for example, when strata correspond to administrative units. Stratified random
sampling can be used to obtain both strata-specific
and population-wide estimates. An equally important motivation for stratified random sampling is
that the logistics of sampling are sometimes simplified, or the cost of sampling is reduced. For example, if a National Forest is to be sampled, it may
be efficient to stratify on a district basis, because
personnel located within each district can sample
that area at a lower cost than centrally located personnel. Cost and feasibility are adequate reasons
for stratified random sampling even if no gain in
accuracy is anticipated. However, it is possible to
degrade the quality of inferential methods if there
are many strata that are not particularly different,
and so it is best not to stratify unless there are
clearly different strata or clear advantages in terms
of feasibility or cost.
For stratified random sampling, the HorvitzThompson estimator is set up as follows. Suppose
that there are Nr popUlation units in the rth stratum. Stratified random sampling specifies that an
SRS of nr units be selected from the rth stratum.
Suppose that the ith population unit is a member
of the rth stratum. Then the inclusion probability
for the ith unit is 7T; = nrlNr. Equation (1) is the
Horvitz-Thompson estimator of the popUlation
total. Stratum-specific estimates are based on
estimates of the stratum total Tr for the rth strata.
The Horvitz-Thompson estimate of Tr is Tr =
