Section 2.3: Neglecting Serial Correlation
15
there was a good negative correlation (Labitzke, 1987; Labitzke and van Loon,
1988).
Labitzke's finding was and is spectacular - and obviously right for the data
from the time interval at her disposal (see Figure 2.1). Of course it could be
that the result was a coincidence as unlikely as the formation of a Mexican
Hat. Or it could represent areal on-going signal. Unfortunately, the data
which were used by Labitzke to formulate her hypothesis can no longer be
used for the assessment of whether we deal with a signal or a coincidence.
Therefore an answer to this question requires information unrelated to the
data as for instance dynamical arguments or GeM experiments. However,
physical hypotheses on the nature ofthe solar-weather link were not available
and are possibly developing right now - so that nothing was left but to wait for
more data and better understanding. (The data which have become available
since Labitzke's discovery in 1987 support the hypothesis.)
In spite of this fundamental problem an intense debate about the "statistical significance" broke out. The reviewers of the first comprehensive paper on
that matter by Labitzke and van Loon (1988) demanded a test. Reluctantly
the authors did wh at they were asked for and found of course an extremely
little risk for the rejection of the null hypothesis "The solar-weather link is
zero". After the publication various other papers were published dealing with
technical aspects of the test - while the basic problem that the data to conduct
the test had been used to formulate the null hypothesis remained.
When hypotheses are to be derived from limited data, I suggest two alternative routes to go. If the time scale of the considered process is short compared
to the available data, then split the full data set into two parts. Derive the
hypothesis (for instance a statistical model) from the first half of the data and
examine the hypothesis with the remaining part of the data. 1 If the time scale
of the considered process is long compared to the time series such that a split
into two parts is impossible, then I recommend using alt data to build a model
optimalty fitting the data. Check the fitted model whether it is consistent with
all known physical features and state explicitly that it is impossible to make
statements about the reliability of the model because of limited evidence.
2.3 N eglecting Serial Correlation
Most standard statistical techniques are derived with explicit need for statistically independent data. However, almost all climatic data are somehow
correlated in time. The resulting problems for testing nullhypotheses is discussed in some detail in Section 9.4. In case of the t-test the problem is
nowadays often acknowlegded - and as a eure people try to determine the
"equivalent sampIe size" (see Section 2.4). When done properly, the t-test
1 An exa.mple of this approach is offered by Wallace and GutzIer (1981).
15
there was a good negative correlation (Labitzke, 1987; Labitzke and van Loon,
1988).
Labitzke's finding was and is spectacular - and obviously right for the data
from the time interval at her disposal (see Figure 2.1). Of course it could be
that the result was a coincidence as unlikely as the formation of a Mexican
Hat. Or it could represent areal on-going signal. Unfortunately, the data
which were used by Labitzke to formulate her hypothesis can no longer be
used for the assessment of whether we deal with a signal or a coincidence.
Therefore an answer to this question requires information unrelated to the
data as for instance dynamical arguments or GeM experiments. However,
physical hypotheses on the nature ofthe solar-weather link were not available
and are possibly developing right now - so that nothing was left but to wait for
more data and better understanding. (The data which have become available
since Labitzke's discovery in 1987 support the hypothesis.)
In spite of this fundamental problem an intense debate about the "statistical significance" broke out. The reviewers of the first comprehensive paper on
that matter by Labitzke and van Loon (1988) demanded a test. Reluctantly
the authors did wh at they were asked for and found of course an extremely
little risk for the rejection of the null hypothesis "The solar-weather link is
zero". After the publication various other papers were published dealing with
technical aspects of the test - while the basic problem that the data to conduct
the test had been used to formulate the null hypothesis remained.
When hypotheses are to be derived from limited data, I suggest two alternative routes to go. If the time scale of the considered process is short compared
to the available data, then split the full data set into two parts. Derive the
hypothesis (for instance a statistical model) from the first half of the data and
examine the hypothesis with the remaining part of the data. 1 If the time scale
of the considered process is long compared to the time series such that a split
into two parts is impossible, then I recommend using alt data to build a model
optimalty fitting the data. Check the fitted model whether it is consistent with
all known physical features and state explicitly that it is impossible to make
statements about the reliability of the model because of limited evidence.
2.3 N eglecting Serial Correlation
Most standard statistical techniques are derived with explicit need for statistically independent data. However, almost all climatic data are somehow
correlated in time. The resulting problems for testing nullhypotheses is discussed in some detail in Section 9.4. In case of the t-test the problem is
nowadays often acknowlegded - and as a eure people try to determine the
"equivalent sampIe size" (see Section 2.4). When done properly, the t-test
1 An exa.mple of this approach is offered by Wallace and GutzIer (1981).
