5.1 Central Limit Theorem and Its Role in Machine Learning
81
Next, we get the following 2 expression of the second term on the right-hand side:
# i
#
−
# i
#
p
2
p
=
# i
#
2 − 2
# i
#
# i
#
p
+
# i
#
2
p
p
=
# i
#
2
p
− 2
# i
#
p
# i
#
p
+
# i
#
2
p
=
# i
#
2
p
−
# i
#
2
p
=
1
#
p i (1 − p i )
,
(5.11)
summarizing all, we find
P
# i
#
− p i
< <
≥ 1 −
p i (1 − p i )
2 #
.
(5.12)
This is what we wanted to show.
Law of large numbers (general case)
So far, we have considered how close the empirical probability is to the true
probability p i . In general, it is known 3 that for an “observable” X, the average of #
trials
X
#
=
1
#
#
n=1
X n
(5.13)
is subject to the following condition:
P
|X
#
− μ| < <
≥ 1 −
σ 2
2 #
,
(5.14)
2 The calculations of the first term can be done in the same way as (5.6):
# i
#
2
p
=
1
#
#
n=1
X
(i)
n
2
p
=
1
# 2
#
n=1,m=1
X
(i)
n X
(i)
m
p
=
1
# 2
#
n=1,m=1
X
(i)
n X
(i)
m
p
=
1
# 2
#
n=1,m=1
p i (m = n)
p 2
i (m = n)
=
1
# 2
#p i + #(# − 1)p
2
i
=
1
#
p i (1 − p i )
+ p
2
i .
(5.10)
The point is that there is a case division in the calculation.
3 The derivation in the general case is left to the reader.
Précédent

- 89/211

Suivant