2.2 Machine Learning and Occam’s Razor
21
Data generation probability known only to God
The above (2.2) is an example of the supervised data {(x[i], d[i])} i=1,2,...# , although
small in size. Supervised data such as MNIST and CIFAR-10 are basically obtained
by repeating the following protocols: 4
1. Choose a label with some probability and express it as d.
2. Take the image corresponding to d by some method and set it as x.
3. Record (x, d).
Then, it is natural to expect that behind all the data in the world there exists a data
generation probability P (x, d) like (2.5), and that the data itself is the result of
sampling according to the probability
(x[i], d[i]) ∼ P (x, d) .
(2.7)
We assume the existence of the data generation probability P (x, d) and
that (2.7) on the data is the starting point for statistical machine learning. It is said
that Einstein wrotes in a letter to Born that “God does not roll the dice,” but machine
learning conversely says “God rolls dice.” Of course, nobody knows the concrete
expression of P (x, d) like (2.5), but it is useful in later discussions to make this
assumption. 5
2.2 Machine Learning and Occam’s Razor
We asserted above that “we do not know the concrete expression of P (x, d).”
However, the goal in machine learning is to “approximately” know the concrete
expression of P (x, d), although they appear to contradict each other. Basically, we
4 This process corresponds to P (x, d) = P (x|d)P (d), but in reality the following order is easier to
collect data:
1. Take an image x.
2. Judge the label of the image and set it to d.
3. Record (x, d).
This process corresponds to P (x, d) = P (d|x)P (x). The resulting sampling should be the same
from Bayes’ theorem (see the column in this chapter), admitting the existence of data generation
probabilities.
5 Needless to say, the probabilities here are all classical, and quantum theory has nothing to do with
it.
Précédent

- 31/211

Suivant