66
4 Advanced Neural Networks
Note that exactly the same J h acts t − 1 times from the right. Since t is the length of
a sentence, it can be very long. Then
|J h | > 1: t (1)| is too large,
(4.23)
|J h | < 1: t (1)| is too small.
(4.24)
So both cases hinder learning. For example, if it is too large, it conflicts with the
“small” updates, which is an implicit assumption that the update expression in (4.20)
is valid. If it is too small, it means that τ = 0 has no effect in the update formula,
which means that the memory is forgotten. These are called exploding gradient
and vanishing gradient, respectively. This was pointed out in Hochreiter’s doctoral
dissertation [44]. 7
4.3 LSTM
As explained above, with a simple recurrent neural network, the gradient explosion/vanishing cannot be improved by using the structure, deepening, convolution,
etc. A new idea is needed to solve this problem. This section describes LSTM (long
short-term memory) [46] which is commonly used currently. 8
Memory vector
The core idea of LSTM is to set up a vector that controls the memory inside the
recurrent neural network. It acts as a RAM (random access memory, temporary
storage area). Let us call the memory vector at time t as
|c(t)
(4.25)
Fig. 4.8 shows the overall diagram of the LSTM.
Here g f , g i , and g o are called
g f : Forget gate
g i : Input gate
g o : Output gate
and we consider them as the sigmoid function for each component,
g f = g i = g o = σ .
(4.26)
7 This is written in German; for English literature see [45].
8 An explanation without the bracket notation is found in [47] on which we largely rely.
Précédent

- 75/211

Suivant