300
F. Firouzi et al.
Fig. 5.53 Local optima in
the training process of neural
network models
Global
Local
• Early stopping, which means the learning is stopped before the overfit occurs
• Dropout technique, in which randomly selected neurons are ignored (dropped
out) during some iterations of the training process. This technique makes the
training process noisy. As a result, neurons take the same level of responsibility
and learn a sparse representation, which makes the model more robust.
Early stopping is an effective method for generalization. In this approach, the
data is divided into three datasets: training set, validation set, and the test set. The
training set is used for the backpropagation algorithm and updating the weights of
the model. During the training process, the error of the validation set is monitored,
which should normally show a decreasing trend. But as the network begins to overfit
the data of the training set, the error of the validation set will increase. This behavior
can be used to stop the training process, and the values of the weights and biases
at the time that the validation set error was minimum can be chosen as the proper
result of the training process. Finally, the test set is used to compare the efficiency
of different models [12, 13].
5.6.3 Activation Function
The activation function is one of the key hyperparameters in a neural network.
Sigmoid (σ (z) =
1
1+exp(−z) ), tanh (tanh(z) =
exp(z)−exp(−z)
exp(z)+exp(−z) ), and ReLU
(ReLU(z) = max (0, z)) are the most commonly used activation functions.
Figure 5.54 represents the plots of these functions [14].
ReLU Basically, what ReLU does is keeping positive input as is while rectifying all
the negative inputs as 0. Accordingly, one of the key advantages of ReLU compared
to other activation functions is that it does not activate all neurons simultaneously by
throwing out all the negative values. This makes it very computationally efficient,
specially when there is a very big and deep neural network consisting of several
layers with dozens of neurons. In practice, ReLU converges much faster than
sigmoid and tanh activation functions.
F. Firouzi et al.
Fig. 5.53 Local optima in
the training process of neural
network models
Global
Local
• Early stopping, which means the learning is stopped before the overfit occurs
• Dropout technique, in which randomly selected neurons are ignored (dropped
out) during some iterations of the training process. This technique makes the
training process noisy. As a result, neurons take the same level of responsibility
and learn a sparse representation, which makes the model more robust.
Early stopping is an effective method for generalization. In this approach, the
data is divided into three datasets: training set, validation set, and the test set. The
training set is used for the backpropagation algorithm and updating the weights of
the model. During the training process, the error of the validation set is monitored,
which should normally show a decreasing trend. But as the network begins to overfit
the data of the training set, the error of the validation set will increase. This behavior
can be used to stop the training process, and the values of the weights and biases
at the time that the validation set error was minimum can be chosen as the proper
result of the training process. Finally, the test set is used to compare the efficiency
of different models [12, 13].
5.6.3 Activation Function
The activation function is one of the key hyperparameters in a neural network.
Sigmoid (σ (z) =
1
1+exp(−z) ), tanh (tanh(z) =
exp(z)−exp(−z)
exp(z)+exp(−z) ), and ReLU
(ReLU(z) = max (0, z)) are the most commonly used activation functions.
Figure 5.54 represents the plots of these functions [14].
ReLU Basically, what ReLU does is keeping positive input as is while rectifying all
the negative inputs as 0. Accordingly, one of the key advantages of ReLU compared
to other activation functions is that it does not activate all neurons simultaneously by
throwing out all the negative values. This makes it very computationally efficient,
specially when there is a very big and deep neural network consisting of several
layers with dozens of neurons. In practice, ReLU converges much faster than
sigmoid and tanh activation functions.
