38
3 Basics of Neural Networks
Here J x , J y , and J are coupling constants. Let us make a model Q J (x, d) using this
“physical system.” First, from (3.6), it is natural to make a Boltzmann distribution
about d, 3
Q J (d|x) =
e −H J,x (d)
Z
=
e (xJ x +yJ y +J )d
1
˜
d=0
e (x μ J μ +J ) ˜
d
=
e (xJ x +yJ y +J )d
1 + e (x μ J μ +J ) .
(3.7)
Here, we introduce a function called the sigmoid function, 4
σ (X) =
1
e −X + 1
.
(3.8)
Then we find
Q J (d = 1|x) = σ (xJ x + yJ y + J ),
(3.9)
Q J (d = 0|x) = 1 − σ (xJ x + yJ y + J ).
(3.10)
This expression is for a given x, but if we suppose the existence of data generation
probability P (x, d), the x should be generated according to
P (x) = P (x, d = 0) + P (x, d = 1) ,
(3.11)
so we use this to define
Q J (x, d) = Q J (d|x)P (x) .
(3.12)
Now that we have a model, let us consider actually adjusting (learning) the
coupling constants J x , J y , and J to mimic the true distribution P (x, d). The relative
entropy (2.8) is calculated from settings such as (3.12),
D KL (P ||Q J ) =
x,d
P (x, d) log
P (d|x)
Q J (d|x)
= −
x,d
P (x, d) log Q J (d|x) + (J -independent part).
(3.13)
3 In the Boltzmann distribution, the factor of H/(k B T ), which is the Hamiltonian divided by the
temperature, is in the exponent. We redefine J to J k B T to absorb the temperature part, so that the
temperature does not appear in the exponent.
4 The reason that the sigmoid function is similar to the Fermi distribution function can be
understood from the fact that d takes on a binary value in the current system. d = 0 corresponds to
the state where the fermion site is not occupied (vacancy state), and d = 1 corresponds to the state
where fermion excitation exists.
Précédent

- 48/211

Suivant