88
Sampling Design and Statistical Inference for Ecological Assessment
'lT23 = 113, because there is only one sample in
which each combination (0 and 2, 0 and 3, and
2 and 3) appear. In general, for an SRS, 'lTij =
n(n - 1)IN(N - 1) (Thompson, 1992); for example, 'lT12 = 2(2 - 1)/(3'2) = 113.
For stratified random sampling, the secondorder inclusion probability 'lTij is n,{nr - 1)IN,{Nr -
1) if the ith and jth popUlation units both belong to
the rth stratum, and 'lTjj = 'lTj'ITj if the ith and jth
population units belong to different strata. Here nr
is the number of units sampled from the rth strata,
and Nr is the number of population units in the rth
strata. For systematic sampling, 'lTij = 'lTj = 11k if
the ith andjth population units both appear in the
sample; otherwise, 'lTij = O. For example, suppose
that every fourth population unit is selected after a
random starting unit has been determined. If population units are exactly 4 units apart, then they are
guaranteed to appear simultaneously in the sample,
and 'lTij = 'lTj = 'lTj = 1/4. If the units are exactly 3
units apart, then they are guaranteed not to appear
in the same sample, and 'lTij = O. Siirndal et al.
(1992) present second-order inclusion probabilities
for other designs.
For any design, Thompson (1992) shows that the
variance of f is estimated by
V(1)= I -~- +2I I - - - YiYj'
~ ~
v (1 -' lTj) v (1 1 )
i=1
'IT,
i=lj>j 'IT(TTj
'lTij
The standard error of f is V vel). The variances of
estimators derived from f are estimated by functions of Ven according to standard properties of
the variance. For instance, the standard error of the
sample mean is estimated by V Ven/v.
6.8 Nonprobability and Gradient
Directed Transect Sampling
Probability sampling is sometimes forsaken in ecological assessment. Often the aim of sampling is to
describe the range and extent of biotic communities across a landscape. In these situations, probability sampling may be inefficient, and sampling
design objectives are oriented toward informal description, rather than formal statistical inference.
Descriptive sampling designs are illustrated by the
sampling protocols used by Cooper et al. (1987),
Pfister et al. (1977), and Steele et al. (1981) to develop the forest habitat type classification systems
for Montana and Idaho. These studies were largely
observational in nature, and they dispensed with
statistical validity in an effort to sample all environments without fully knowing the range of environments and plant communities before sampling
commenced. The authors employed a sampling
procedure termed by Mueller-Dombois and Ellenburg (1974) as "subjective, but without preconceived bias" to sample forest stands across large areas of remote and rugged mountains. Sampling was
conducted by traveling forest roads and sampling
whenever a forest stand was encountered that appeared representative of the local forest community. In contrast to the sampling design used by
these authors, simple random sampling would be
expected to sample representatively with respect to
the areal coverage of forest communities. If an SRS
had been obtained, unless sample sizes were substantially larger than those used by these authors,
sample sizes may have been inadequate to describe
some communities with limited distribution.
Gradient directed transect, or gradsect, sampling
(Austin and Heyligers, 1989; Bourgeron et aI.,
1994) is a more useful nonprobability sampling design than Mueller-Dombois and Ellenburg'S (1974)
approach. Two objectives of gradsect sampling are
to obtain complete coverage of the range of environmental conditions and biotic communities and
to obtain information for constructing probability
sampling designs. Gradsect sampling is used to
sample regional-scale landscapes in a manner that
distributes plots representatively with respect to
major environmental gradients. This is accomplished by selecting sample locations to maximize
variation between units with respect to major recognized environmental variables, such as elevation
and moisture availability (Bourgeron et aI., 1994).
An assumption made of gradsect sampling is that
the sample is representative of the landscape, and
the validity of this assumption depends on the ability of the researcher to recognize the variables controlling the distribution of ecological communities.
Because extrapolation, inference, and prediction
hinge on the validity of this assumption, inference
should be made with caution and accompanied by
a clear statement of the conditional nature of the
inference.
In some instances, probability sampling designs
fail to accurately identify the population of interest.
Also, researchers sometimes fail to sample the population according to design because of logistical
problems that occur during the sampling process. A
practical approach to these problems is to make a
best effort at selecting a probability sample and admit that (1) the resultant estimators are nearly unbiased, rather than unbiased; (2) predictions are
moderately precise, rather than highly precise; and
(3) hypothesis tests nearly, rather than exactly,
Sampling Design and Statistical Inference for Ecological Assessment
'lT23 = 113, because there is only one sample in
which each combination (0 and 2, 0 and 3, and
2 and 3) appear. In general, for an SRS, 'lTij =
n(n - 1)IN(N - 1) (Thompson, 1992); for example, 'lT12 = 2(2 - 1)/(3'2) = 113.
For stratified random sampling, the secondorder inclusion probability 'lTij is n,{nr - 1)IN,{Nr -
1) if the ith and jth popUlation units both belong to
the rth stratum, and 'lTjj = 'lTj'ITj if the ith and jth
population units belong to different strata. Here nr
is the number of units sampled from the rth strata,
and Nr is the number of population units in the rth
strata. For systematic sampling, 'lTij = 'lTj = 11k if
the ith andjth population units both appear in the
sample; otherwise, 'lTij = O. For example, suppose
that every fourth population unit is selected after a
random starting unit has been determined. If population units are exactly 4 units apart, then they are
guaranteed to appear simultaneously in the sample,
and 'lTij = 'lTj = 'lTj = 1/4. If the units are exactly 3
units apart, then they are guaranteed not to appear
in the same sample, and 'lTij = O. Siirndal et al.
(1992) present second-order inclusion probabilities
for other designs.
For any design, Thompson (1992) shows that the
variance of f is estimated by
V(1)= I -~- +2I I - - - YiYj'
~ ~
v (1 -' lTj) v (1 1 )
i=1
'IT,
i=lj>j 'IT(TTj
'lTij
The standard error of f is V vel). The variances of
estimators derived from f are estimated by functions of Ven according to standard properties of
the variance. For instance, the standard error of the
sample mean is estimated by V Ven/v.
6.8 Nonprobability and Gradient
Directed Transect Sampling
Probability sampling is sometimes forsaken in ecological assessment. Often the aim of sampling is to
describe the range and extent of biotic communities across a landscape. In these situations, probability sampling may be inefficient, and sampling
design objectives are oriented toward informal description, rather than formal statistical inference.
Descriptive sampling designs are illustrated by the
sampling protocols used by Cooper et al. (1987),
Pfister et al. (1977), and Steele et al. (1981) to develop the forest habitat type classification systems
for Montana and Idaho. These studies were largely
observational in nature, and they dispensed with
statistical validity in an effort to sample all environments without fully knowing the range of environments and plant communities before sampling
commenced. The authors employed a sampling
procedure termed by Mueller-Dombois and Ellenburg (1974) as "subjective, but without preconceived bias" to sample forest stands across large areas of remote and rugged mountains. Sampling was
conducted by traveling forest roads and sampling
whenever a forest stand was encountered that appeared representative of the local forest community. In contrast to the sampling design used by
these authors, simple random sampling would be
expected to sample representatively with respect to
the areal coverage of forest communities. If an SRS
had been obtained, unless sample sizes were substantially larger than those used by these authors,
sample sizes may have been inadequate to describe
some communities with limited distribution.
Gradient directed transect, or gradsect, sampling
(Austin and Heyligers, 1989; Bourgeron et aI.,
1994) is a more useful nonprobability sampling design than Mueller-Dombois and Ellenburg'S (1974)
approach. Two objectives of gradsect sampling are
to obtain complete coverage of the range of environmental conditions and biotic communities and
to obtain information for constructing probability
sampling designs. Gradsect sampling is used to
sample regional-scale landscapes in a manner that
distributes plots representatively with respect to
major environmental gradients. This is accomplished by selecting sample locations to maximize
variation between units with respect to major recognized environmental variables, such as elevation
and moisture availability (Bourgeron et aI., 1994).
An assumption made of gradsect sampling is that
the sample is representative of the landscape, and
the validity of this assumption depends on the ability of the researcher to recognize the variables controlling the distribution of ecological communities.
Because extrapolation, inference, and prediction
hinge on the validity of this assumption, inference
should be made with caution and accompanied by
a clear statement of the conditional nature of the
inference.
In some instances, probability sampling designs
fail to accurately identify the population of interest.
Also, researchers sometimes fail to sample the population according to design because of logistical
problems that occur during the sampling process. A
practical approach to these problems is to make a
best effort at selecting a probability sample and admit that (1) the resultant estimators are nearly unbiased, rather than unbiased; (2) predictions are
moderately precise, rather than highly precise; and
(3) hypothesis tests nearly, rather than exactly,
