378
A. Jmal et al.
3 Long Short Term Memory-RNN/Back Propagation
Neural Networks
Recurrent Neural Networks are the only networks with internal memory, which
makes them robust and powerful. In a RNN, Weights are applied to both the
current input and the looping back output and are adjusted through gradient
descent or back propagation [12]. The RNN work on this recursive formula (1)
where X t is the input at time step t, S t is the state at time step t and F w is the
recursive function.
S t = F w ∗ (S t−1 , X t )
(1)
S t = F w (S t−1 , X t )
(2)
S t = tanh(W s ∗ S t−1 , W x ∗ X t )
(3)
Y t = W y ∗ S t
(4)
The recursive function is a tanh function (3), we multiply the input state with
the weights of X mentioned as W x and the previous state with W s and then past
it through a tanh activation to get the new state (3). To get the output vector,
we multiply the new state S t with W y (4). RNN learn use back propagation
through time. Therefore, we calculate the loss using the output, go back to
each state, and update weights by multiplying gradients. The updating weights
would be negligible and our network will not get any better. This problem is
called vanishing gradients problem. To solve it and to improve the accuracy,
we add a more interactions to RNN and this is the idea behind Long Short
Term Memory (LSTM) [13]. The LSTM cell is capable of learning long-term
dependencies. RNNs usually have a short memory and are extended by LSTM
units to extend the memory of the network. It provides the capabilities to absorb
more information from even longer sequences of data. This helps to boost the
precision of the prediction by taking into account more data. The LSTM cell
maintains three kinds of gates and one cell state: the input gate, the forget gate
and the output gate. The architecture of an LSTM cell is shown in Fig. 2, where
the input gate chooses what new information needs to be stored in the cell state.
This is shown in Eq. 5 and 7, where i t is the input gate layer output and C t is the
cell state update. The forget gate decides what existing information in cell state
needs to be thrown away, this is shown in Eq. 8, where C t is again the update
of the cell state [12]. Finally, the output gate filters the output and determines
the final cell output. This can be seen through Eqs. 9 and 10, where o t is the
output-gate layer output and h t is the resulting hidden state for the given input.
ˇ
C t is called as intermediate cell state, used to calculate the C t , which is the cell
state using Eq. 6. The input gate and the intermediate cell state are added with
the old cell state and the forget gate, and then this cell state is passed through
tanh activation to be multiplied with the output gate [14]. The following steps
are used to train a LSTM-RNN [14]:
A. Jmal et al.
3 Long Short Term Memory-RNN/Back Propagation
Neural Networks
Recurrent Neural Networks are the only networks with internal memory, which
makes them robust and powerful. In a RNN, Weights are applied to both the
current input and the looping back output and are adjusted through gradient
descent or back propagation [12]. The RNN work on this recursive formula (1)
where X t is the input at time step t, S t is the state at time step t and F w is the
recursive function.
S t = F w ∗ (S t−1 , X t )
(1)
S t = F w (S t−1 , X t )
(2)
S t = tanh(W s ∗ S t−1 , W x ∗ X t )
(3)
Y t = W y ∗ S t
(4)
The recursive function is a tanh function (3), we multiply the input state with
the weights of X mentioned as W x and the previous state with W s and then past
it through a tanh activation to get the new state (3). To get the output vector,
we multiply the new state S t with W y (4). RNN learn use back propagation
through time. Therefore, we calculate the loss using the output, go back to
each state, and update weights by multiplying gradients. The updating weights
would be negligible and our network will not get any better. This problem is
called vanishing gradients problem. To solve it and to improve the accuracy,
we add a more interactions to RNN and this is the idea behind Long Short
Term Memory (LSTM) [13]. The LSTM cell is capable of learning long-term
dependencies. RNNs usually have a short memory and are extended by LSTM
units to extend the memory of the network. It provides the capabilities to absorb
more information from even longer sequences of data. This helps to boost the
precision of the prediction by taking into account more data. The LSTM cell
maintains three kinds of gates and one cell state: the input gate, the forget gate
and the output gate. The architecture of an LSTM cell is shown in Fig. 2, where
the input gate chooses what new information needs to be stored in the cell state.
This is shown in Eq. 5 and 7, where i t is the input gate layer output and C t is the
cell state update. The forget gate decides what existing information in cell state
needs to be thrown away, this is shown in Eq. 8, where C t is again the update
of the cell state [12]. Finally, the output gate filters the output and determines
the final cell output. This can be seen through Eqs. 9 and 10, where o t is the
output-gate layer output and h t is the resulting hidden state for the given input.
ˇ
C t is called as intermediate cell state, used to calculate the C t , which is the cell
state using Eq. 6. The input gate and the intermediate cell state are added with
the old cell state and the forget gate, and then this cell state is passed through
tanh activation to be multiplied with the output gate [14]. The following steps
are used to train a LSTM-RNN [14]:
