6.5 Hypothesis Testing and p values
The accuracy of an estimator is a measure of
closeness between the estimates and the parameter
of interest. Accuracy reflects both bias and precision, because an estimator may be precise without
being unbiased. An accurate estimator has low bias
and high precision. Accuracy is usually defined
mathematically to be the mean squared error (MSE)
of the estimator, which is the expected value of the
mean squared difference between the estimator and
the parameter of interest; that is, MSE(S) = E[(S -
0 2 ], where g is the parameter of interest. Mathematically, the MSE is the sum of the variance
(]"~ and the squared bias; that is, MSE(S) = (]"~ +
[E(S) - g]2.
6.4.2 Confidence Intervals
It is rarely sufficient to provide an estimate without an accompanying measure of its reliability. For
example, an estimate of the standard error of an estimator provides a measure of reliability. Confidence intervals are often more useful, however, because they are sets of candidate values for the
parameter that are in agreement with the sample
data. Confidence intervals are constructed so that
there is a high probability of bracketing, or capturing, the true value of the parameter between the
lower and upper limits. To illustrate the interpretation of confidence intervals, suppose that an estimate fr of the proportion of infected stands 1T in a
watershed is 0.29. A hypothetical 95% confidence
interval for the true population proportion is [0.23,
0.35]. Thus, if the same sampling design were employed to sample the same population infinitely often and confidence intervals constructed from each
sample, then 95% of these intervals would contain
the true population proportion, no matter what the
true value might be.
To clarify the interpretation of confidence intervals, suppose that 1T denotes the unknown population proportion. Let L(y) and V(y) denote the lower
and upper confidence limits computed from the
random sample Y. Note that these limits are statistics and hence random variables. The confidence
limits are constructed from Y so that the probability that the interval covers 1T is 0.95. Mathematically, this probability is expressed as 0.95 =
Pr[L(Y) < 1T < V(y)], What is sometimes confusing about confidence intervals is that the statement
0.95 = Pr[0.23 < 1T < 0.35] is false, because 1T is
not random and neither are the numbers 0.23 and
0.35. The statement 0.23 < 1T < 0.35 is either absolutely true or absolutely false; there can be no
probabilistic interpretation in classical frequentist
statistics. The correct interpretation is that the
83
method used to construct the interval will bracket
1T about 95% of the time, and not that there is a
0.95 probability that 1T is between 0.23 and 0.35.
A good sampling design will produce confidence
intervals with narrow widths and achieve the stated
confidence level.
6.5 Hypothesis Testing and
p Values
Occasionally, the objective of sampling is to formally test a hypothesis about a population. Hypothesis testing is appropriate when the researcher
identifies a theory, postulate, or hypothesis to be
established or confirmed before the sample is collected. Testing a hypothesis that is suggested by a
sample by using the sample itself greatly increases
the risk of drawing an incorrect conclusion from
the test.
A hypothesis test requires a statement of null and
alternative hypotheses. The alternative hypothesis
usually is the hypothesis of interest, and the null
hypothesis represents the contrary to the alternative. For instance, suppose that it is thought that
white pine blister rust is becoming more prevalent
in a watershed, and the historical stand infection
rate is believed to be 0.2 or less. If 1T represents the
current unknown stand infection rate, then a useful
null hypothesis is Ho: 1T = 0.2, and an appropriate
alternative is HI: 1T> 0.2. To test the null hypothesis, suppose that an SRS of stands is selected and
the sample proportion of infected stands is observed to be 0.29. This value supports HI and contradicts Ho, but the strength of support is not obvious. A test statistic is a function of the sample that
is used to measure how strongly the data contradict Ho. Strong evidence against Ho lends credence
to the contrary, as expressed by HI. If the evidence
is strong enough, then we reject Ho and conclude
that HI is true. If the evidence is not strong enough,
then we conclude that the evidence is not strong
enough to conclude that HI is true. The second conclusion is weak, because the hypothesis test is intended only to measure evidence against Ho; it does
not measure the strength of evidence supporting Ho.
Furthermore, we began with a preconceived notion
that Ho was false, and we are reluctant to give up
the notion without evidence supporting its truth.
The p value is a quantitative measure of the
strength of evidence against the null hypothesis.
For example, if Z(y) is a test statistic, such as a Z
statistic computed from the sample proportion (see
Ostle and Mensing, 1975), then large values of the
random variable Z(y) = (fr - 1T)/frir correspond to
Précédent

- 93/539

Suivant