20
2 Introduction to Machine Learning
(x[1] = 6, d[1] = 1) (x[10] = 6, d[10] = 1) (x[19] = 6, d[19] = 1)
(x[2] = 6, d[2] = 1) (x[11] = 3, d[11] = 0) (x[20] = 6, d[20] = 1)
(x[3] = 6, d[3] = 1) (x[12] = 3, d[12] = 0) (x[21] = 6, d[21] = 1)
(x[4] = 6, d[4] = 1) (x[13] = 6, d[13] = 1) (x[22] = 6, d[22] = 1)
(x[5] = 6, d[5] = 0) (x[14] = 4, d[14] = 0) (x[23] = 6, d[23] = 1)
(x[6] = 6, d[6] = 1) (x[15] = 3, d[15] = 0) (x[24] = 5, d[24] = 0)
(x[7] = 6, d[7] = 0) (x[16] = 6, d[16] = 0) (x[25] = 5, d[25] = 0)
(x[8] = 2, d[8] = 0) (x[17] = 1, d[17] = 0) (x[26] = 6, d[26] = 1)
(x[9] = 5, d[9] = 0) (x[18] = 6, d[18] = 1) (x[27] = 4, d[27] = 0)
(2.2)
Now we define a probability with “ˆ” (hat):
ˆ
P (x, d) =
Number of times (x, d) appeared in the data
Total number of data
.
(2.3)
An example experiment taking 1000 data calculates the value of ˆ
P (x, d) as follows:
ˆ
P (x, d) =
d = 0 d = 1
x = 1 0.086 0.0
x = 2 0.087 0.0
x = 3 0.074 0.0
x = 4 0.089 0.0
x = 5 0.083 0.0
x = 6 0.082 0.499
(2.4)
The more data we have, the closer our table is to
P (x, d) =
d = 0 d = 1
x = 1 1/12 0
x = 2 1/12 0
x = 3 1/12 0
x = 4 1/12 0
x = 5 1/12 0
x = 6 1/12 1/2
, (1/12 = 0.08 ˙
3, 1/2 = 0.5).
(2.5)
In fact, the definition shows that (2.5) is the realization probability of (x, d). The
data in (2.2) can be regarded as a sampling from P (x, d). Let us write it as follows:
(x[i], d[i]) ∼ P (x, d) .
(2.6)
2 Introduction to Machine Learning
(x[1] = 6, d[1] = 1) (x[10] = 6, d[10] = 1) (x[19] = 6, d[19] = 1)
(x[2] = 6, d[2] = 1) (x[11] = 3, d[11] = 0) (x[20] = 6, d[20] = 1)
(x[3] = 6, d[3] = 1) (x[12] = 3, d[12] = 0) (x[21] = 6, d[21] = 1)
(x[4] = 6, d[4] = 1) (x[13] = 6, d[13] = 1) (x[22] = 6, d[22] = 1)
(x[5] = 6, d[5] = 0) (x[14] = 4, d[14] = 0) (x[23] = 6, d[23] = 1)
(x[6] = 6, d[6] = 1) (x[15] = 3, d[15] = 0) (x[24] = 5, d[24] = 0)
(x[7] = 6, d[7] = 0) (x[16] = 6, d[16] = 0) (x[25] = 5, d[25] = 0)
(x[8] = 2, d[8] = 0) (x[17] = 1, d[17] = 0) (x[26] = 6, d[26] = 1)
(x[9] = 5, d[9] = 0) (x[18] = 6, d[18] = 1) (x[27] = 4, d[27] = 0)
(2.2)
Now we define a probability with “ˆ” (hat):
ˆ
P (x, d) =
Number of times (x, d) appeared in the data
Total number of data
.
(2.3)
An example experiment taking 1000 data calculates the value of ˆ
P (x, d) as follows:
ˆ
P (x, d) =
d = 0 d = 1
x = 1 0.086 0.0
x = 2 0.087 0.0
x = 3 0.074 0.0
x = 4 0.089 0.0
x = 5 0.083 0.0
x = 6 0.082 0.499
(2.4)
The more data we have, the closer our table is to
P (x, d) =
d = 0 d = 1
x = 1 1/12 0
x = 2 1/12 0
x = 3 1/12 0
x = 4 1/12 0
x = 5 1/12 0
x = 6 1/12 1/2
, (1/12 = 0.08 ˙
3, 1/2 = 0.5).
(2.5)
In fact, the definition shows that (2.5) is the realization probability of (x, d). The
data in (2.2) can be regarded as a sampling from P (x, d). Let us write it as follows:
(x[i], d[i]) ∼ P (x, d) .
(2.6)
