40
3 Basics of Neural Networks
d 1
d 4
d 3
d 2
x
y
J
Fig. 3.2 Left : Data with (1, 0, 0, 0), blue; (0, 1, 0, 0), red; (0, 0, 1, 0), green; (0, 0, 0, 1), black.
Right : Schematic diagram of the Hamiltonian (3.18)
system, but the deep neural network which will be introduced later has a slightly
more complicated structure, so if there is no good gradient calculation method,
calculation time will be consumed . Fortunately, there is an effective way to calculate
gradients derived from the network structure. This will be explained later.
Multiclass classification
Suppose that the teaching signal is d = (d 1 , d 2 , d 3 , d 4 ) with
4
I =1 d I = 1, and let
d 1 be blue, d 2 be red, d 3 be green, d 4 be black as shown in Fig. 3.2. Since d has
increased to 4 components, the model Hamiltonian is modified to
H J,x (d) = −
4
I =1
(xJ xI + yJ yI + J I )d
I .
(3.18)
Since it is troublesome to write the sum symbol explicitly, we shall use a matrix
notation: J = (J xI , J yI ) (4 × 2 matrix), J = (J 1 , J 2 , J 3 , J 4 ) (4-dimensional vector),
x = (x, y) (2-dimensional vector). Then,
H J,x (d) = −d · (Jx + J) .
(3.19)
The Boltzmann weight with this Hamiltonian is
Q J (d|x) =
e −H J,x (d)
Z
=
e d·(Jx+J)
e (Jx+J) 1 + e (Jx+J) 2 + e (Jx+J) 3 + e (Jx+J) 4
.
(3.20)
Précédent

- 50/211

Suivant