22
Chapter 2: Misuses
icance level by comparing the value of (2.12) with the percentiles of the
t( n* )-distribution.
.
The advantage of (2.12) is that this test operates as specified by the user
provided that the interval between successive observations is long enough.
The disadvantage is that a reduced amount of data is utilized in the analysis. Therefore, the following concept was developed in the 1970s to overcome
this disadvantage: The numerator in (2.12) is a random variable because it
differs from one pair of temperature sampies to the next. When the observations which comprise the sampies are serially uncorrelated the denominator
in (2.12) is an estimate of the standard deviation of the numerator and the
ratio can be thought of as an expression of the difference of means in units
of estimated standard deviations. For serially correlated data, with sampie
means T and sampie standard deviations (1 derived from all available observations, the standard deviation of TH - Tv is .)«(1k + (1~ )/n' with the
equivalent sampie size n' as defined in (2.6). For sufficiently large sampies
sizes the ratio
TH-TV
t = ---r;=;;==::::;;=::::;:=
.)(o-k + o-~ )/n'
(2.13)
has a standard Gaussian distribution with zero mean and standard deviation
one. Thus one can conduct a test by comparing (2.13) to the percentiles of
the standard Gaussian distribution.
So far everything is fine.
Since t( n') is approximately equal to the Gaussian distribution for n' ~ 30,
one may compare the test statistic (2.13) also with the percentiles of the
t( n')-distribution. The incorrect step is the heuristic assumption that this
prescription - "compare with the percentiles of the t(n'), or t(n' - 1) distribution" - would be right for sm all (n' < 30) equivalent sampies sizes. The
rationale of doing so is the tacitly assumed fact that the statistic (2.13) would
be t(n') or t(n' - l)-distributed under the null hypothesis. However, this assumption is simply wrong. The distribution (2.13) is not t(k)-distributed for
any k, be it the equivalent sampie size n' or any other number. This result
has been published by several authors (Katz (1982), Thiebaux and Zwiers
(1984) and Zwiers and von Storch (1994)) but has stubbornly been ignored
by most of the atmospheric sciences community.
A justification for the small sampie size test would be that its behaviour
under the null hypothesis is weIl approximated by the t-test with the equivalent sampie size representing the degrees of freedom. But this is not so, as
is demonstrated by the following example with an AR(l)-process (2.3) with
a = .60. The exact equivalent sampie size n' = ~n is known for the process
since its parameters are compietely known. One hundred independent sampIes of variable length n were randomly generated. Each sampie was used to
test the null hypothesis Ho : E(Xt ) = 0 with the t-statistic (2.13) at the 5%
significance level. If the test operates correctly the null hypothesis should be
(incorrectly) rejected 5% of the time. The actual rejection rate (Figure 2.4)
Chapter 2: Misuses
icance level by comparing the value of (2.12) with the percentiles of the
t( n* )-distribution.
.
The advantage of (2.12) is that this test operates as specified by the user
provided that the interval between successive observations is long enough.
The disadvantage is that a reduced amount of data is utilized in the analysis. Therefore, the following concept was developed in the 1970s to overcome
this disadvantage: The numerator in (2.12) is a random variable because it
differs from one pair of temperature sampies to the next. When the observations which comprise the sampies are serially uncorrelated the denominator
in (2.12) is an estimate of the standard deviation of the numerator and the
ratio can be thought of as an expression of the difference of means in units
of estimated standard deviations. For serially correlated data, with sampie
means T and sampie standard deviations (1 derived from all available observations, the standard deviation of TH - Tv is .)«(1k + (1~ )/n' with the
equivalent sampie size n' as defined in (2.6). For sufficiently large sampies
sizes the ratio
TH-TV
t = ---r;=;;==::::;;=::::;:=
.)(o-k + o-~ )/n'
(2.13)
has a standard Gaussian distribution with zero mean and standard deviation
one. Thus one can conduct a test by comparing (2.13) to the percentiles of
the standard Gaussian distribution.
So far everything is fine.
Since t( n') is approximately equal to the Gaussian distribution for n' ~ 30,
one may compare the test statistic (2.13) also with the percentiles of the
t( n')-distribution. The incorrect step is the heuristic assumption that this
prescription - "compare with the percentiles of the t(n'), or t(n' - 1) distribution" - would be right for sm all (n' < 30) equivalent sampies sizes. The
rationale of doing so is the tacitly assumed fact that the statistic (2.13) would
be t(n') or t(n' - l)-distributed under the null hypothesis. However, this assumption is simply wrong. The distribution (2.13) is not t(k)-distributed for
any k, be it the equivalent sampie size n' or any other number. This result
has been published by several authors (Katz (1982), Thiebaux and Zwiers
(1984) and Zwiers and von Storch (1994)) but has stubbornly been ignored
by most of the atmospheric sciences community.
A justification for the small sampie size test would be that its behaviour
under the null hypothesis is weIl approximated by the t-test with the equivalent sampie size representing the degrees of freedom. But this is not so, as
is demonstrated by the following example with an AR(l)-process (2.3) with
a = .60. The exact equivalent sampie size n' = ~n is known for the process
since its parameters are compietely known. One hundred independent sampIes of variable length n were randomly generated. Each sampie was used to
test the null hypothesis Ho : E(Xt ) = 0 with the t-statistic (2.13) at the 5%
significance level. If the test operates correctly the null hypothesis should be
(incorrectly) rejected 5% of the time. The actual rejection rate (Figure 2.4)
