7.6 Hypothesis Testing and p-Values
85
For an example let us return to the question posed in the introduction to this chapter
whether all fit parameters are really necessary. Therefore we consider fitting data to
a straight line as shown in Fig. 7.1. It appears reasonable to fit the linear dependence
y i = a + bs i to the dataset but we might even consider a second-order polynomial
y i = a + bs i + cs
2
i , which will follow the data points even closer and result in a
smaller χ
2 , because we have the additional parameter c to approximate the dataset.
But since we expect the data to lie on the straight line we state the hypothesis that
c = 0. We test it by fitting the second-order polynomial to the dataset, then determine
the error bars σ (c) of the fit parameter c using (7.7), and finally calculate the test
statistics t = c/σ (c). If we have many degrees of freedom ν = n − m 1, we can
approximate ν (t) by a Gaussian and check whether t is larger than 2, which would
indicate that our hypothesis c = 0 is rejected at the 10% level. Conversely, if t is
smaller, we corroborate the hypothesis that c is consistent with zero and we might
as well omit it from the fit.
The discussed method works well to test hypotheses about individual fit parameters, but occasionally we have to figure out whether we can omit a larger number of
fit parameters at the same time. This is the topic of the following section.
7.7 F-Test
In the previous section we used the t-statistic to determine whether one coefficient
in a regression model is compatible with zero and can be omitted. This works well if
there is only a single obsolete coefficient, but might fail if several of the coefficients
are significantly correlated. In that case the error bars for each of the coefficients
is large and leads to the conclusion that the coefficient can be omitted, but the real
origin of the problem is that the fitting procedure fails to work out whether to assign
the uncertainty to one or the other(s) of the correlated coefficients. This leads to
serious mis-interpretations of fit results.
One way to solve the dilemma with several potentially correlated coefficients
and to determine which ones to omit is the F-test. In this test the contribution of
a group of p fit parameters to the χ
2
p is determined. We define this by the squared
difference of the N “measurement” values y i and the regression model, normalized
to the measurement error σ i , thus s i = (y i −
p
j=1 A i j x j )/σ i . For χ
2
p we then obtain
χ
2
p =
N
i=1
s
2
i =
N
i=1
y i −
p
j=1 A i j x j
σ i
2
.
(7.36)
In the next step we increase the number of fit parameters to q > p and test whether χ
2
q
is significantly smaller than χ
2
p . In order to quantify this improvement, we introduce
the F-statistic
85
For an example let us return to the question posed in the introduction to this chapter
whether all fit parameters are really necessary. Therefore we consider fitting data to
a straight line as shown in Fig. 7.1. It appears reasonable to fit the linear dependence
y i = a + bs i to the dataset but we might even consider a second-order polynomial
y i = a + bs i + cs
2
i , which will follow the data points even closer and result in a
smaller χ
2 , because we have the additional parameter c to approximate the dataset.
But since we expect the data to lie on the straight line we state the hypothesis that
c = 0. We test it by fitting the second-order polynomial to the dataset, then determine
the error bars σ (c) of the fit parameter c using (7.7), and finally calculate the test
statistics t = c/σ (c). If we have many degrees of freedom ν = n − m 1, we can
approximate ν (t) by a Gaussian and check whether t is larger than 2, which would
indicate that our hypothesis c = 0 is rejected at the 10% level. Conversely, if t is
smaller, we corroborate the hypothesis that c is consistent with zero and we might
as well omit it from the fit.
The discussed method works well to test hypotheses about individual fit parameters, but occasionally we have to figure out whether we can omit a larger number of
fit parameters at the same time. This is the topic of the following section.
7.7 F-Test
In the previous section we used the t-statistic to determine whether one coefficient
in a regression model is compatible with zero and can be omitted. This works well if
there is only a single obsolete coefficient, but might fail if several of the coefficients
are significantly correlated. In that case the error bars for each of the coefficients
is large and leads to the conclusion that the coefficient can be omitted, but the real
origin of the problem is that the fitting procedure fails to work out whether to assign
the uncertainty to one or the other(s) of the correlated coefficients. This leads to
serious mis-interpretations of fit results.
One way to solve the dilemma with several potentially correlated coefficients
and to determine which ones to omit is the F-test. In this test the contribution of
a group of p fit parameters to the χ
2
p is determined. We define this by the squared
difference of the N “measurement” values y i and the regression model, normalized
to the measurement error σ i , thus s i = (y i −
p
j=1 A i j x j )/σ i . For χ
2
p we then obtain
χ
2
p =
N
i=1
s
2
i =
N
i=1
y i −
p
j=1 A i j x j
σ i
2
.
(7.36)
In the next step we increase the number of fit parameters to q > p and test whether χ
2
q
is significantly smaller than χ
2
p . In order to quantify this improvement, we introduce
the F-statistic
