6.10 Example of a Large-scale Sampling Design
achieve a specified significance level. The difficulties of probability sampling for ecosystem assessment do not absolve researchers of the responsibility of striving for probability samples.
Researchers must be knowledgeable about and use
probability sampling when feasible, and be able to
recognize the difference between a probability and
a nonprobability sample. When probability sampling cannot be achieved, researchers should attempt to devise a sampling design that comes as
close as possible to selecting a probability sample.
This strategy defends against estimation bias and
promotes efficiency in terms of the amount of information recovered from the sample for a given
cost. The researcher is responsible for revealing the
difficulties and failure of the design to his or her
audience so that they may interpret the results objectively. Scientifically literate audiences usually
are willing to consider an analysis based on nonprobability samples if provided with adequate explanations of the failure to obtain probability samples and if statistical inference is limited and
conservative.
6.9 Advanced Sampling
Techniques
In addition to the designs discussed earlier, advanced sampling techniques can be used to increase
sampling efficiency. For example, adaptive sampling designs (Thompson, 1992) allow samplers to
deviate from a sampling plan when a rare or unusual population unit is observed. Adaptive sampling allows the researcher to modify the sample
design, while sampling is in progress, to increase
the chances of observing another population unit
with similar characteristics. It also provides a means
for computing inclusion probabilities for all sample
units. Consequently, the researcher is able to improve sampling efficiency and collect a statistically
valid sample. Double sampling (Schreuder et al.,
1993) is a method that reduces sampling costs by
sampling a surrogate, or auxiliary, variable. Both
the variable of interest and the surrogate are measured on some sample units; on the remainder, only
the surrogate is measured at a reduced cost. Then
sample units on which both are measured are used
to predict the variable of interest for the surrogateonly sample units. For example, double sampling
is sometimes used when estimating animal population size using aerial surveys. Ground visitation
of a subsample of the aerial sample locations is
used to correct the aerial only observations for undercounting (Thompson, 1992).
6.10 Example of a Large-scale
Sampling Design
89
The National Surface Water Survey (NSWS) was
initiated during the 1980s by the U.S. Environmental Protection Agency to investigate acidification of surface waters. The National Stream Survey (NSS) was charged with investigating the water
chemistry of streams sensitive to acidic deposition.
Objectives of the NSS were to detect and monitor
regional patterns and trends in surface-water acidity and to analyze the relationship between observed patterns and trends in surface-water quality
and atmospheric deposition (Ward et aI., 1990).
The population of interest was streams in southeastern and Mid-Atlantic states, and the population
units were defined to be stream reaches appearing
on I :250,000-scale topographic maps (Kaufmann
et aI., 1988). A stream reach was defined as "a segment of stream between two confluences, or the
segment between the origin of the stream and the
first confluence if the stream was a headwaters
reach" (Overton and Stehman, 1995, p. 261). Because of the scale and cost of the survey, a pilot
study was used to assess the feasibility of the proposed sampling design (Ward et aI., 1990). The pilot study sampled 54 stream reaches, and the data
were used in conjunction with Monte Carlo simulation to compare and select statistical methods for
trend detection of low-concentration chemicals.
The full sample design stratified the study area
by region and subregion. Sampling units were selected by randomly locating a lattice over a subregion map. If a lattice point fell on a stream reach,
then the stream reach was designated to be sampled (Overton and Stehman, 1995). There are several consequences of this design. Inclusion probabilities were proportional to stream reach area and
unknown until the sample was selected. Once the
sample units were selected into the sample, the inelusion probabilities for each sample unit were
computed based on the number of lattice points, the
area of each sample unit (a stream reach), and the
total area of the subregion (Kaufmann et aI., 1988).
The sampling design did not use a sampling frame
and hence avoided the cost and effort of constructing a sampling frame. (Recall that the purpose of the sampling frame is to list each population unit and its inclusion probability so that each
unit can be sampled according to its inclusion probability. The use of the sampling frame in data analysis is restricted to identifying the inclusion probabilities for the sample units.) The sampling design
yielded 445 stream reaches for sampling. Both the
achieve a specified significance level. The difficulties of probability sampling for ecosystem assessment do not absolve researchers of the responsibility of striving for probability samples.
Researchers must be knowledgeable about and use
probability sampling when feasible, and be able to
recognize the difference between a probability and
a nonprobability sample. When probability sampling cannot be achieved, researchers should attempt to devise a sampling design that comes as
close as possible to selecting a probability sample.
This strategy defends against estimation bias and
promotes efficiency in terms of the amount of information recovered from the sample for a given
cost. The researcher is responsible for revealing the
difficulties and failure of the design to his or her
audience so that they may interpret the results objectively. Scientifically literate audiences usually
are willing to consider an analysis based on nonprobability samples if provided with adequate explanations of the failure to obtain probability samples and if statistical inference is limited and
conservative.
6.9 Advanced Sampling
Techniques
In addition to the designs discussed earlier, advanced sampling techniques can be used to increase
sampling efficiency. For example, adaptive sampling designs (Thompson, 1992) allow samplers to
deviate from a sampling plan when a rare or unusual population unit is observed. Adaptive sampling allows the researcher to modify the sample
design, while sampling is in progress, to increase
the chances of observing another population unit
with similar characteristics. It also provides a means
for computing inclusion probabilities for all sample
units. Consequently, the researcher is able to improve sampling efficiency and collect a statistically
valid sample. Double sampling (Schreuder et al.,
1993) is a method that reduces sampling costs by
sampling a surrogate, or auxiliary, variable. Both
the variable of interest and the surrogate are measured on some sample units; on the remainder, only
the surrogate is measured at a reduced cost. Then
sample units on which both are measured are used
to predict the variable of interest for the surrogateonly sample units. For example, double sampling
is sometimes used when estimating animal population size using aerial surveys. Ground visitation
of a subsample of the aerial sample locations is
used to correct the aerial only observations for undercounting (Thompson, 1992).
6.10 Example of a Large-scale
Sampling Design
89
The National Surface Water Survey (NSWS) was
initiated during the 1980s by the U.S. Environmental Protection Agency to investigate acidification of surface waters. The National Stream Survey (NSS) was charged with investigating the water
chemistry of streams sensitive to acidic deposition.
Objectives of the NSS were to detect and monitor
regional patterns and trends in surface-water acidity and to analyze the relationship between observed patterns and trends in surface-water quality
and atmospheric deposition (Ward et aI., 1990).
The population of interest was streams in southeastern and Mid-Atlantic states, and the population
units were defined to be stream reaches appearing
on I :250,000-scale topographic maps (Kaufmann
et aI., 1988). A stream reach was defined as "a segment of stream between two confluences, or the
segment between the origin of the stream and the
first confluence if the stream was a headwaters
reach" (Overton and Stehman, 1995, p. 261). Because of the scale and cost of the survey, a pilot
study was used to assess the feasibility of the proposed sampling design (Ward et aI., 1990). The pilot study sampled 54 stream reaches, and the data
were used in conjunction with Monte Carlo simulation to compare and select statistical methods for
trend detection of low-concentration chemicals.
The full sample design stratified the study area
by region and subregion. Sampling units were selected by randomly locating a lattice over a subregion map. If a lattice point fell on a stream reach,
then the stream reach was designated to be sampled (Overton and Stehman, 1995). There are several consequences of this design. Inclusion probabilities were proportional to stream reach area and
unknown until the sample was selected. Once the
sample units were selected into the sample, the inelusion probabilities for each sample unit were
computed based on the number of lattice points, the
area of each sample unit (a stream reach), and the
total area of the subregion (Kaufmann et aI., 1988).
The sampling design did not use a sampling frame
and hence avoided the cost and effort of constructing a sampling frame. (Recall that the purpose of the sampling frame is to list each population unit and its inclusion probability so that each
unit can be sampled according to its inclusion probability. The use of the sampling frame in data analysis is restricted to identifying the inclusion probabilities for the sample units.) The sampling design
yielded 445 stream reaches for sampling. Both the
