4.3 LSTM
69
Note that since the weight from the forget gate and the input from the input gate
are placed at every τ , the information recorded once is not completely retained. In
the memory vector, the “forget” and “remember” operations controlled only by J are
given in advance in an explicit form. For example, let us take the following memory
vector:
|c =
⎛
⎜
⎜
⎜
⎝
c 1
c 2
. . .
c ♠
⎞
⎟
⎟
⎟
⎠
(4.36)
and assume that c 1 = (the component that governs a subject). If the input is I
enjoy machine learning., and a successfully trained LSTM is a machine
that “predicts the next word,” when the first word I comes in, the memory should
be able to work as
|c =
⎛
⎜
⎜
⎜
⎝
1
0
. . .
0
⎞
⎟
⎟
⎟
⎠
(4.37)
and recognize the situation as “we are dealing with the subject part no”. 9 At this
stage, it may be followed by and you ..., so it is not possible to cancel the
“subject” state. When the second word enjoy is input, since this is a verb, the
subject will no longer follow in English grammar. In this case, the output of the
forget gate is
|g f =
⎛
⎜
⎜
⎜
⎝
0
1
. . .
1
⎞
⎟
⎟
⎟
⎠
,
(4.38)
and when |c is multiplied as with this, it is an operation to make c 1 = (the
component that controls the subject) zero. Further, if the second component of the
memory vector = (the component that controls the verb), the input gate when the
second word enjoy is input is, for example,
9 Keep in mind that this is just for illustration. In actual situations everything is dealt with in real
numbers, and there are many kinds of different subjects. Here, consider the setting for convenience
of explanation.
Précédent

- 78/211

Suivant