5 Machine Learning for IoT
301
Leaky ReLU
ReLU
Tanh
Sigmoid
0
0
0
1
-1
0
0
1
Fig. 5.54 Most popular activation functions
Sigmoid In practical applications, the sigmoid activation function is less used
despite its popularity in the past because of two important problems:
1. Sigmoid function eliminates the gradient, which means when the input value is
very big or very small, the output of the function is either 1 or 0. This will result
in a vanishing gradient problem in the backpropagation algorithm. In this case,
no signal is transmitted through the neurons and thus the neuron will not learn
anything in the training phase.
2. The outputs of the sigmoid function are not zero centered, and in the backpropagation algorithm, this will create gradients that are either all positive or
all negative. This is also not appropriate for the gradient updates of the weights.
Tanh Similarly, the tanh activation function has the vanishing gradient problem.
However, in contrast to the second problem of the sigmoid function, its output is
zero centered.
ReLU ReLU, an abbreviation of rectified linear unit, has gained great popularity in
recent years. It has been reported that ReLU activation function greatly impacts
the convergence rate of stochastic gradient descent algorithms, six times more
than tanh and sigmoid functions [6–15]. The primary reason for this phenomenon
is its linearity, which causes the gradient not to vanish. Besides, ReLU is less
computationally expensive because it is simple thresholding at zero compared to
tanh and sigmoid functions that depend on exponential and complex operations.
Note that ReLU is sparsely activated because it is zero for all negative inputs. This
sparsity (i.e., not all the neurons are active at the same time) might be a good thing
because it can reduce the power of the neural network, resulting in less overfitting.
The downside is that it can also lead to dying ReLU problem. A dead ReLU refers to
the case where a ReLU always generates outputs with the same value (zero) which
is not important for any inputs and next neuron layers, resulting in an incomplete
learning process and very weak mode. ReLU indeed has another problem as well.
The output of neurons with ReLU activation function can grow dramatically because
the output of ReLU is a linear function of its input (for positive values). In other
words, ReLU does not have any kind of boundary, and thus it cannot truncate
the output. Compare it with sigmoid or tanh functions, in which the outputs are
saturated.
Précédent

- 307/647

Suivant