6.4 Statistical Methods
successively selecting units from the population so
that every unit has the same probability of being
selected on each draw. On the first draw, the probability that a particular unit is selected is liN; on
the second draw, the probability that a particular
unit (not selected on the first draw) is selected is
l/(N - 1), and so on. For example, all possible
samples of size 2 that can be selected from the population P = {a, 2, 3} are {a, 2}, {a, 3}, and {2, 3}.
Each sample must have a probability of t of being
selected in order that the sampling design be an
SRS. Techniques for random selection of samples
are discussed by Levy and Lemeshow (1991) and
Thompson (1992).
6.2.3 Experimental Design
What is the difference between experimental and
sampling designs? If a treatment is purposively imposed on the objects or organisms of interest, then
the study and its design are experimental. In experimental studies, a probability sample is chosen
from a population, and some of the sampling units
are manipulated with the intent of affecting or eliciting a response from the units. Sampling designs
do not manipulate sample units with this intent. For
example, an experiment to test whether the presence of a suspected disease host (e.g., gooseberry)
increases the infection rate of a pathogen (e.g.,
white pine blister rust) selects a probability sample
from a population of stands. Then, some fraction
of selected stands is randomly chosen to be cleared
of gooseberry, and the remaining stands are left untreated. Measurement of the proportion of infected
trees before and after clearing the stands provides
data for a test of the hypothesis that removal of
gooseberry does not affect the population infection
rate, versus an alternative hypothesis stating that
removal of gooseberry decreases the population infection rate. In this case, some of the sample units
(i.e., stands) have been purposely manipulated by
the researcher. In contrast, a sample design aimed
at answering the same question collects a probability sample of stands, but no stands are cleared.
An informative statistic that may be computed from
the sample data is the sample correlation coefficient, measuring the degree of association between
the proportion of infected trees and the areal cover
of gooseberry. Both studies may provide convincing evidence that the rate of infection is positively
associated with gooseberry coverage, but only the
experiment provides the opportunity to obtain statistical evidence supporting the conclusion that
gooseberry is a cause of infection in the population. The sampling design cannot rule out the possibility that gooseberry does not cause infection,
81
but instead responds positively to an environmental variable that also favors the pathogen. For more
information on statistical methods for environmental studies, see Eberhart and Thomas (1991) for a
useful overview. Also, Petersen (1994) is a recent
text on experimental design oriented toward biologists.
6.3 Random Sampling
The idea of a random sample must be understood
to discuss statistical inference in more detail. A
random sample is not a subset of population units,
but a process by which a sample is selected from
a population and a variable or attribute is observed
on the sampled units. The nature of the process is
random in the sense that the outcome of the sampling process is not known ahead of time. However, the distribution of possible outcomes of the
process is determined by the population and the
sampling design.
Let Y = (Ylo ... ,Yn) denote a random sample.
Once a sample is collected and the observations are
recorded, then we have a realization of Y, and the
realized sample is denoted by y = (Yl, ... , Yn). For
example, Y may represent the process of selecting
an SRS of size 2 from the population P = {a, 2,
3}. A possible realization of Y is Y = {a, 2}. A statistic is a function of the random sample whose realizations are numerical. Because it is a function of
a random process, a statistic is also a random variable. For instance, the sample mea..!!. is a random
variable, and it may be expressed as X(y), alt~ugh
the usual expression for the sample mean is Y.
Suppose that every possible sample is selected
from a population, and a statistic S is computed using every one of the samples. The expected value
of S, denoted by E(S), is the weighted average of
all its possible values; the weight associated with
each value is the probability that the value is the
realized value. For the population P = {a, 2, 3},
the population mean is J.L = t. An SRS specifies
that each of the three subsets of size 2 listed in Section 6.2.2 has a probability of t of being chosen.
The possible realizations of the sample mean are 1,
1.5, and 2.5; each is equally likely. Hence the expected value of the sample mean is (1 X 113) +
(1.5 X 113) + (2.5 X 113) = 5/3.
6.4 Statistical Methods
Two rather different views of the population arise
in sample design and statistical analysis. The first,
and traditional, view is design based; it treats the
successively selecting units from the population so
that every unit has the same probability of being
selected on each draw. On the first draw, the probability that a particular unit is selected is liN; on
the second draw, the probability that a particular
unit (not selected on the first draw) is selected is
l/(N - 1), and so on. For example, all possible
samples of size 2 that can be selected from the population P = {a, 2, 3} are {a, 2}, {a, 3}, and {2, 3}.
Each sample must have a probability of t of being
selected in order that the sampling design be an
SRS. Techniques for random selection of samples
are discussed by Levy and Lemeshow (1991) and
Thompson (1992).
6.2.3 Experimental Design
What is the difference between experimental and
sampling designs? If a treatment is purposively imposed on the objects or organisms of interest, then
the study and its design are experimental. In experimental studies, a probability sample is chosen
from a population, and some of the sampling units
are manipulated with the intent of affecting or eliciting a response from the units. Sampling designs
do not manipulate sample units with this intent. For
example, an experiment to test whether the presence of a suspected disease host (e.g., gooseberry)
increases the infection rate of a pathogen (e.g.,
white pine blister rust) selects a probability sample
from a population of stands. Then, some fraction
of selected stands is randomly chosen to be cleared
of gooseberry, and the remaining stands are left untreated. Measurement of the proportion of infected
trees before and after clearing the stands provides
data for a test of the hypothesis that removal of
gooseberry does not affect the population infection
rate, versus an alternative hypothesis stating that
removal of gooseberry decreases the population infection rate. In this case, some of the sample units
(i.e., stands) have been purposely manipulated by
the researcher. In contrast, a sample design aimed
at answering the same question collects a probability sample of stands, but no stands are cleared.
An informative statistic that may be computed from
the sample data is the sample correlation coefficient, measuring the degree of association between
the proportion of infected trees and the areal cover
of gooseberry. Both studies may provide convincing evidence that the rate of infection is positively
associated with gooseberry coverage, but only the
experiment provides the opportunity to obtain statistical evidence supporting the conclusion that
gooseberry is a cause of infection in the population. The sampling design cannot rule out the possibility that gooseberry does not cause infection,
81
but instead responds positively to an environmental variable that also favors the pathogen. For more
information on statistical methods for environmental studies, see Eberhart and Thomas (1991) for a
useful overview. Also, Petersen (1994) is a recent
text on experimental design oriented toward biologists.
6.3 Random Sampling
The idea of a random sample must be understood
to discuss statistical inference in more detail. A
random sample is not a subset of population units,
but a process by which a sample is selected from
a population and a variable or attribute is observed
on the sampled units. The nature of the process is
random in the sense that the outcome of the sampling process is not known ahead of time. However, the distribution of possible outcomes of the
process is determined by the population and the
sampling design.
Let Y = (Ylo ... ,Yn) denote a random sample.
Once a sample is collected and the observations are
recorded, then we have a realization of Y, and the
realized sample is denoted by y = (Yl, ... , Yn). For
example, Y may represent the process of selecting
an SRS of size 2 from the population P = {a, 2,
3}. A possible realization of Y is Y = {a, 2}. A statistic is a function of the random sample whose realizations are numerical. Because it is a function of
a random process, a statistic is also a random variable. For instance, the sample mea..!!. is a random
variable, and it may be expressed as X(y), alt~ugh
the usual expression for the sample mean is Y.
Suppose that every possible sample is selected
from a population, and a statistic S is computed using every one of the samples. The expected value
of S, denoted by E(S), is the weighted average of
all its possible values; the weight associated with
each value is the probability that the value is the
realized value. For the population P = {a, 2, 3},
the population mean is J.L = t. An SRS specifies
that each of the three subsets of size 2 listed in Section 6.2.2 has a probability of t of being chosen.
The possible realizations of the sample mean are 1,
1.5, and 2.5; each is equally likely. Hence the expected value of the sample mean is (1 X 113) +
(1.5 X 113) + (2.5 X 113) = 5/3.
6.4 Statistical Methods
Two rather different views of the population arise
in sample design and statistical analysis. The first,
and traditional, view is design based; it treats the
