80
7 Regression Models and Hypothesis Testing
N (x; μ, σ ) =
1
√
2πσ
exp
−
(x − μ)
2
2σ 2
.
(7.25)
Note that ¯
X n and S
2
n are commonly called the sample average and the unbiased
sample variance, respectively.
2 We first need to understand how fast the estimates
¯
X n and S
2
n converge towards the true values μ and σ
2 as we increase n. We therefore
calculate the expectation value of the squared difference between the estimate and
the true value
¯
X n − μ
2
=
1
n
i
x i − μ
2
=
1
n
i
(x i − μ)
2
=
1
n 2
i
j
(x i − μ)(x j − μ)
(7.26)
=
1
n 2
i
(x i − μ)
2
=
1
n 2
i
σ
2
=
σ
2
n
,
where we exploited the fact that the different x i are statistically independent and
therefore (x i − μ)(x j − μ) = =(x i − μ)
2
δ i j = σ
2
δ i j . Finally we see that the estimate ¯
X n on average approaches the true mean μ with σ/
√
n as we increase the
number of samples n.
Next we need to address the reliability of an estimate based on a small number
n of samples. To achieve this goal we introduce a test-statistic t, which is given by
the deviation of the estimate X n from the real mean μ, but divided by the estimated
standard deviation S n /
√
n.
t =
¯
X n − μ
S n /
√
n
=
( ¯
X n − μ)/(σ/
√
n)
S 2
n /σ 2
.
(7.27)
In the second equality in (7.27) we divide numerator and denominator by σ to
visualize that the numerator x = ( ¯
X n − μ)/(σ/
√
n) indeed stems from a normal
distribution with unit variance, given by
ψ x (x) =
1
√
2π
e
−x
2 /2
.
(7.28)
Moreover, the denominator y =
S 2
n /σ 2 stems from the square root of a χ
2 -
distribution, the latter given by (7.22). Thus t = x/y is a random variable, defined by
the ratio of a Gaussian random variable x and a random variable y, which is derived
2 Defining the sample variance with n in the denominator yields a value that is too small and is
called biassed, because using ¯
X n , which is derived from the same samples as S 2
n , is closer to the
samples x i than the “true” mean μ. This is compensated by dividing by n − 1 instead, which leads
to the unbiased sample variance, defined in (7.24).
7 Regression Models and Hypothesis Testing
N (x; μ, σ ) =
1
√
2πσ
exp
−
(x − μ)
2
2σ 2
.
(7.25)
Note that ¯
X n and S
2
n are commonly called the sample average and the unbiased
sample variance, respectively.
2 We first need to understand how fast the estimates
¯
X n and S
2
n converge towards the true values μ and σ
2 as we increase n. We therefore
calculate the expectation value of the squared difference between the estimate and
the true value
¯
X n − μ
2
=
1
n
i
x i − μ
2
=
1
n
i
(x i − μ)
2
=
1
n 2
i
j
(x i − μ)(x j − μ)
(7.26)
=
1
n 2
i
(x i − μ)
2
=
1
n 2
i
σ
2
=
σ
2
n
,
where we exploited the fact that the different x i are statistically independent and
therefore (x i − μ)(x j − μ) = =(x i − μ)
2
δ i j = σ
2
δ i j . Finally we see that the estimate ¯
X n on average approaches the true mean μ with σ/
√
n as we increase the
number of samples n.
Next we need to address the reliability of an estimate based on a small number
n of samples. To achieve this goal we introduce a test-statistic t, which is given by
the deviation of the estimate X n from the real mean μ, but divided by the estimated
standard deviation S n /
√
n.
t =
¯
X n − μ
S n /
√
n
=
( ¯
X n − μ)/(σ/
√
n)
S 2
n /σ 2
.
(7.27)
In the second equality in (7.27) we divide numerator and denominator by σ to
visualize that the numerator x = ( ¯
X n − μ)/(σ/
√
n) indeed stems from a normal
distribution with unit variance, given by
ψ x (x) =
1
√
2π
e
−x
2 /2
.
(7.28)
Moreover, the denominator y =
S 2
n /σ 2 stems from the square root of a χ
2 -
distribution, the latter given by (7.22). Thus t = x/y is a random variable, defined by
the ratio of a Gaussian random variable x and a random variable y, which is derived
2 Defining the sample variance with n in the denominator yields a value that is too small and is
called biassed, because using ¯
X n , which is derived from the same samples as S 2
n , is closer to the
samples x i than the “true” mean μ. This is compensated by dividing by n − 1 instead, which leads
to the unbiased sample variance, defined in (7.24).
