2.2 Machine Learning and Occam’s Razor
21
Data generation probability known only to God
The above (2.2) is an example of the supervised data {(x[i], d[i])} i=1,2,...# , although
small in size. Supervised data such as MNIST and CIFAR-10 are basically obtained
by repeating the following protocols: 4
1. Choose a label with some probability and express it as d.
2. Take the image corresponding to d by some method and set it as x.
3. Record (x, d).
Then, it is natural to expect that behind all the data in the world there exists a data
generation probability P (x, d) like (2.5), and that the data itself is the result of
sampling according to the probability
(x[i], d[i]) ∼ P (x, d) .
(2.7)
We assume the existence of the data generation probability P (x, d) and
that (2.7) on the data is the starting point for statistical machine learning. It is said
that Einstein wrotes in a letter to Born that “God does not roll the dice,” but machine
learning conversely says “God rolls dice.” Of course, nobody knows the concrete
expression of P (x, d) like (2.5), but it is useful in later discussions to make this
assumption. 5
2.2 Machine Learning and Occam’s Razor
We asserted above that “we do not know the concrete expression of P (x, d).”
However, the goal in machine learning is to “approximately” know the concrete
expression of P (x, d), although they appear to contradict each other. Basically, we
4 This process corresponds to P (x, d) = P (x|d)P (d), but in reality the following order is easier to
collect data:
1. Take an image x.
2. Judge the label of the image and set it to d.
3. Record (x, d).
This process corresponds to P (x, d) = P (d|x)P (x). The resulting sampling should be the same
from Bayes’ theorem (see the column in this chapter), admitting the existence of data generation
probabilities.
5 Needless to say, the probabilities here are all classical, and quantum theory has nothing to do with
it.
21
Data generation probability known only to God
The above (2.2) is an example of the supervised data {(x[i], d[i])} i=1,2,...# , although
small in size. Supervised data such as MNIST and CIFAR-10 are basically obtained
by repeating the following protocols: 4
1. Choose a label with some probability and express it as d.
2. Take the image corresponding to d by some method and set it as x.
3. Record (x, d).
Then, it is natural to expect that behind all the data in the world there exists a data
generation probability P (x, d) like (2.5), and that the data itself is the result of
sampling according to the probability
(x[i], d[i]) ∼ P (x, d) .
(2.7)
We assume the existence of the data generation probability P (x, d) and
that (2.7) on the data is the starting point for statistical machine learning. It is said
that Einstein wrotes in a letter to Born that “God does not roll the dice,” but machine
learning conversely says “God rolls dice.” Of course, nobody knows the concrete
expression of P (x, d) like (2.5), but it is useful in later discussions to make this
assumption. 5
2.2 Machine Learning and Occam’s Razor
We asserted above that “we do not know the concrete expression of P (x, d).”
However, the goal in machine learning is to “approximately” know the concrete
expression of P (x, d), although they appear to contradict each other. Basically, we
4 This process corresponds to P (x, d) = P (x|d)P (d), but in reality the following order is easier to
collect data:
1. Take an image x.
2. Judge the label of the image and set it to d.
3. Record (x, d).
This process corresponds to P (x, d) = P (d|x)P (x). The resulting sampling should be the same
from Bayes’ theorem (see the column in this chapter), admitting the existence of data generation
probabilities.
5 Needless to say, the probabilities here are all classical, and quantum theory has nothing to do with
it.
