84
Sampling Design and Statistical Inference for Ecological Assessment
large sample proportions and contradict Ho. Suppose that the observed datay yields a value of 1.96
when used to compute Z; that is, Z(y) = 1.96. The
approximate normality of Z(y) implies that the p
value is approximately Pr[Z(Y) ;::: 1.96] = 0.025,
which says that the probability is approximately
0.025 that we would have observed a Z statistic as
large or larger than 1.96 given that 7r were truly
0.2. The p value is the probability of obtaining as
much or more evidence contradicting Ho and favoring HI. given Ho were true. Hence a p value as
small as 0.025 indicates that either Ho is not true
or an unlikely sample has been collected from a
population with an infection rate of 0.2. Because
we began with the notion that Ho is false, we state
that the data contradict Ho and support HI, but acknowledge the possibility that we may be wrong.
These ideas can be formalized in terms of a decision rule. We reject Ho when the p value is less
than or equal to a predetermined significance level.
A popular choice of significance level is 0.05, although there are often good reasons for using different values. The significance level is the probability that the test will incorrectly reject Ho when
Ho is true. Rejecting Ho when Ho is true is called
a Type I error. Hence the significance level is the
probability of a Type I error. We may also fail to
reject Ho when Ho is false and thus commit a Type
II error. Type I errors are controlled by the significance level of the test. For example, if the significance level is chosen to be 0.0001, there is very
little chance of committing a Type I error. Unfortunately, this strategy of minimizing Type I errors
increases the risk of Type II errors. Type II errors
are controlled by choosing an efficient sample design and ensuring that the sample size is adequate.
The use of hypothesis testing has drawn much
criticism. Although the mathematics of hypothesis
testing are elegant and the methodology works well
in simple situations, the logic of hypothesis testing
leaves much to be desired, as do the difficulties of
interpreting multiple tests. In many instances, confidence interval construction can be used to accomplish the same general objectives as hypothesis testing. Hypothesis testing, confidence interval
construction, and the relationships between the two
are extensively discussed in applied statistics texts.
Useful references are Ramsey and Schafer (1997),
Ostle and Mensing (1975), and Snedecor and
Cochran (1980).
6.5.1 Sample Size
A central topic in sampling design is that of choosing the sample size. Sample size strongly determines the accuracy of estimates and predictions,
along with the Type II error rate of hypothesis tests.
For example, suppose that an objective of the design is to estimate the population mean J.L by calculating a point estimate and a confidence interval.
The width of the confidence interval can be controlled in the sample design by controlling the size
of the standard error of the estimator of J.L. For an
SRS design, the variance of the sample mean is
up = (N - n)u 2 /(Nn), where u2 is the population
variance. An approximate 90% confidence interval
for the pop~lation mean has lower an~ upper limits L(Y) = Y - 1.645uy and V(y) = Y + 1.645uy
(Snedecor and Cochran, 1980). The width of the
interval is w = 3.29uy = 3.29uV (N - n)/(Nn),
and if we can provide an educated guess of the
value of u, then we can estimate the minimal sample size by solving for n.
Sample size calculations as simple and exact as
the last example are rarely attainable for more complicated problems, such as estimating the relationship between a response variable and a set of explanatory variables. Approximations are sometimes
possible. For example, the sample size for the regression problem can be approximated by calculating the sample size for a simpler problem, such
as estimating the response variable mean using the
sample mean. Provided that the explanatory variables are truly informative with respect to modeling the response variable, then the variance of the
sample mean will be larger than the variance of the
regression estimator of the response variable mean,
given a set of explanatory variables. Then the estimated sample size for the simpler problem will
be adequate for estimating the regression model.
An alternative approach is to use Monte Carlo
methods to simulate sample data and the distribution of the statistics of interest. Examples of sample size calculations for ecological monitoring are
given by Kendall et al. (1992) and Lesica and Steele
(1996).
6.6 Basic Sampling Designs
A unifying concept in sampling design is the
Horvitz-Thompson estimator (Horvitz and Thompson, 1952; Overton and Stehman, 1995; Thompson, 1992). The Horvitz-Thompson estimator f
gives an unbiased estimate of the population total
T for any probability sampling design. In this section, we use the Horvitz-Thompson estimator to
develop connections between the basic sampling
designs and estimators of T and other population
parameters. Estimators of many other parameters
Précédent

- 94/539

Suivant