84
5 Sampling
Meaning of central limit theorem in machine learning
Let us go back to the first example of the law of large numbers. When the true
probability is not given while only (5.2) is given, the most likely estimation q i of
the true probability is
q i =
# i
#
(5.20)
(see footnote 5 in Chap. 1). According to the central limit theorem, this “best”
estimation q i is subject to
With approximately 70% possibility, p i −
p i (1 − p i )
#
< q i < p i +
p i (1 − p i )
#
.
(5.21)
When regarding (5.20) as the result of machine learning, (5.21) is considered to
represent the generalization performance of this “machine.” In order to have the
generalization, that is, to make q i very close to p i , it is better to take a larger sample
number #. This is reminiscent of the generalized performance inequality (2.11)
using the VC dimensions described in Chap. 2. Actually, the
√
# in the denominator
of the second term on the right-hand side of the inequality (2.11) of generalization
performance is the same as (5.21), and it suggests that the part comes from
the central limit theorem. Thus, the central limit theorem is closely related to
generalization performance.
5.2 Various Sampling Methods
Next, let us look at models. In this book, the model of supervised machine learning
is a conditional probability with parameters,
Q J (d|x) .
(5.22)
When we actually “operate” this model, we want to label some input according to
its probability. Here, there exists a gap between knowing the value of the probability
and sampling according to that probability. For example, if the reader wants to create
a machine that outputs one of 1, 2, 3, 4, 5, 6 with a probability of 1/6, then what
does she/he do? The easiest solution is to make a dice, and for the sampling, just roll
the dice. However, in the model (5.22), the probability of d varies for each input x.
It is ridiculous to keep making dice of various shapes each time.
To make matters worse, there are machine learning models that make it more
difficult to calculate the value of the probability in the first place, as described in the
Précédent

- 92/211

Suivant