4.2 Recurrent Neural Network and Backpropagation
65
we find
δd(t)|h(t) =
τ ≤t
t (τ )|δJ x |x(τ ) + +δ t (τ )|δJ h |h(τ − 1)
=
τ ≤t
m,n
t (τ )|m )δJ
mn
x + +δ t (τ )|m − 1)δJ
mn
h
=
m,n
τ ≤t
t (τ )|m )δJ
mn
x +
τ ≤t
t (τ )|m − 1)δJ
mn
h
.
(4.18)
Therefore,
δL =
t
δd(t)|h(t)
=
m,n
t
τ ≤t
t (τ )|m )δJ
mn
x +
t
τ ≤t
t (τ )|m − 1)δJ
mn
h
,
(4.19)
and then, the update rule is
δJ
mn
x = −
t
τ ≤t
t (τ )|m ), δJ
mn
h = −
t
τ ≤t
t (τ )|m − 1) .
(4.20)
Also,
t (τ − 1)| = =δ t (τ )|J h G(τ − 1), t (t)| = =d(t)|G(t)
(4.21)
is the formula for backpropagation in the present case.
Exploding gradient/vanishing gradient
Using the method explained above, we can train a recurrent neural network with
data, but is that enough? For example, can we train the network using language
data to make a good sentence? The answer is no. For example, it cannot handle
parenthesis structure well. That is, loss of memory occurs. In the following, we
explain why this happens in terms of the backpropagation equation.
Since we adjust the parameter J according to (4.20), let us start from those
equations. t (τ )| commonly appears in both update expressions. Considering the
τ = 0 state, from the backpropagation formula (4.21) we find
t (1)| = =δ t (2)|J h G(1) = =δ t (3)|J h G(2)J h G(1) = . . .
= =d(t)|G(t)J h G(t − 1) . . . J h G(2)J h G(1).
(4.22)
65
we find
δd(t)|h(t) =
τ ≤t
t (τ )|δJ x |x(τ ) + +δ t (τ )|δJ h |h(τ − 1)
=
τ ≤t
m,n
t (τ )|m )δJ
mn
x + +δ t (τ )|m − 1)δJ
mn
h
=
m,n
τ ≤t
t (τ )|m )δJ
mn
x +
τ ≤t
t (τ )|m − 1)δJ
mn
h
.
(4.18)
Therefore,
δL =
t
δd(t)|h(t)
=
m,n
t
τ ≤t
t (τ )|m )δJ
mn
x +
t
τ ≤t
t (τ )|m − 1)δJ
mn
h
,
(4.19)
and then, the update rule is
δJ
mn
x = −
t
τ ≤t
t (τ )|m ), δJ
mn
h = −
t
τ ≤t
t (τ )|m − 1) .
(4.20)
Also,
t (τ − 1)| = =δ t (τ )|J h G(τ − 1), t (t)| = =d(t)|G(t)
(4.21)
is the formula for backpropagation in the present case.
Exploding gradient/vanishing gradient
Using the method explained above, we can train a recurrent neural network with
data, but is that enough? For example, can we train the network using language
data to make a good sentence? The answer is no. For example, it cannot handle
parenthesis structure well. That is, loss of memory occurs. In the following, we
explain why this happens in terms of the backpropagation equation.
Since we adjust the parameter J according to (4.20), let us start from those
equations. t (τ )| commonly appears in both update expressions. Considering the
τ = 0 state, from the backpropagation formula (4.21) we find
t (1)| = =δ t (2)|J h G(1) = =δ t (3)|J h G(2)J h G(1) = . . .
= =d(t)|G(t)J h G(t − 1) . . . J h G(2)J h G(1).
(4.22)
