68
4 Advanced Neural Networks
First, the first half (A) is transformed as follows:
(A) = δ|g o = δ
m
|m o
o
x |x(t) + J
o
h |h(t − 1)
=
m
|m
o (•)
o
x |x(t) + J
o
h |h(t − 1)
=
m
|m
o (•)
o
x |x(t) + δJ
o
h |h(t − 1) + J
o
h δ|h(t − 1)
.
(4.33)
The third item is a form in which the parameter J
o
h acts on δ|h one time step
before, so if we go back more and more time, we have a lot of J
o
h , causing the
exploding/vanishing gradient problem. Next, the second half (B) is written as
(B) = δ
m
|m tanh(
=
m
|m tanh
(•)
δ|c(t)
(C)
,
(4.34)
so the magnitude of the gradient is left to the behavior of (C).
(C) = δ
m
|m − 1) f
+ δ
m
|m tanh
x |x(t) + J h |h(t − 1)
i
=
m
|m
− 1) f + +m|c(t − 1) δ|g f
(D)
+
m
|m
tanh
(•)
x |x(t) + δJ h |h(t − 1) + J h δ|h(t − 1)
i
+ tanh(•) δ|g i
(F )
.
(4.35)
Since (D) and (F) are basically the same as g o , J f and J i act many times, possibly
resulting in the exploding/vanishing gradient problem. However, in the first term
of (4.35), |c(t − 1) does not have J. This part is the backpropagation to the
memory vector one time step ago, and plays the role of “keeping the memory as
much as possible”—since the learning parameter J is not acting, in principle, it can
propagate as long as it can. This is actually clear from the definition: |x(t) |h(t)
is always input to the activation function, whereas |c(t) is just multiplied by some
constant by g f . This is the most important part of the LSTM, the point where the
exploding/vanishing gradient problem is less likely to occur.
4 Advanced Neural Networks
First, the first half (A) is transformed as follows:
(A) = δ|g o = δ
m
|m o
o
x |x(t) + J
o
h |h(t − 1)
=
m
|m
o (•)
o
x |x(t) + J
o
h |h(t − 1)
=
m
|m
o (•)
o
x |x(t) + δJ
o
h |h(t − 1) + J
o
h δ|h(t − 1)
.
(4.33)
The third item is a form in which the parameter J
o
h acts on δ|h one time step
before, so if we go back more and more time, we have a lot of J
o
h , causing the
exploding/vanishing gradient problem. Next, the second half (B) is written as
(B) = δ
m
|m tanh(
=
m
|m tanh
(•)
δ|c(t)
(C)
,
(4.34)
so the magnitude of the gradient is left to the behavior of (C).
(C) = δ
m
|m − 1) f
+ δ
m
|m tanh
x |x(t) + J h |h(t − 1)
i
=
m
|m
− 1) f + +m|c(t − 1) δ|g f
(D)
+
m
|m
tanh
(•)
x |x(t) + δJ h |h(t − 1) + J h δ|h(t − 1)
i
+ tanh(•) δ|g i
(F )
.
(4.35)
Since (D) and (F) are basically the same as g o , J f and J i act many times, possibly
resulting in the exploding/vanishing gradient problem. However, in the first term
of (4.35), |c(t − 1) does not have J. This part is the backpropagation to the
memory vector one time step ago, and plays the role of “keeping the memory as
much as possible”—since the learning parameter J is not acting, in principle, it can
propagate as long as it can. This is actually clear from the definition: |x(t) |h(t)
is always input to the activation function, whereas |c(t) is just multiplied by some
constant by g f . This is the most important part of the LSTM, the point where the
exploding/vanishing gradient problem is less likely to occur.
