78
5 Sampling
then we may find
Q J (grandfather|x) =
1
2
, Q J (dog|x) =
1
2
.
Of course, this is fine, but if we sell this machine to our customers, one of them
would complain as, 1
Should x be classified as a grandfather or a dog?!
In such a case, one irresponsible solution is to use a coin toss:
d = grandfather, if the coin shows the front
d = dog, if the coin shows the back
(5.1)
In this way we can respond to that complaint. Of course, based on (5.1), if the
same image is used to ask the question many times, it will be judged sometimes as
“grandfather” and sometimes as “dog.” It will be no problem because showing the
image to humans will result in the same.
The same problem occurs when actually creating the training data {(x[i], d[i])} i ,
even before using the model Q J which was trained. That is, when many objects
a, b are shown in a single data item x[i], the problem is whether the correct answer
should be d[i] = a, or d[i] = b. Again, this depends on the judgment of the person
creating the data, and that information is included in the data generation probability
P (x, d).
Although we have been silent until now about this issue, from the above
considerations, we can see that it is necessary to consider what it means to actually
sample when there is a probability distribution P . So, in this chapter, we explain the
basics about and around the sampling, by focusing on important items related to:
• sampling of training data (data collection)
• sampling from trained machines (data imitation)
5.1 Central Limit Theorem and Its Role in Machine
Learning
Let us recall the example given in Chap. 1 for the introduction to machine learning.
We have possible events A 1 , A 2 , . . . , A W , and the probability of occurrence is
p 1 , p 2 , . . . , p W , but we do not know the specific value of the probability p i . Instead,
1 This is, of course, a joke to illustrate the need for sampling.
Précédent

- 86/211

Suivant