Intrusion Detection Based on Convolutional Neural Network in Complex . . .
233
data sets. During model training, some hidden nodes are actively and temporarily discarded with a probability p, thus, the network size is greatly reduced. CNN
is allowed to learn features in these incomplete networks, so as to ensure better
generalization ability and reduce the possibility of overfitting effectively.
In the prediction, the results of all sub-models are averaged to increase
the capacity and generalization of the model. The authors in [9,10] insist that
dropout works best and generates the most abundant network structure when
p = 0.5.
Activation Functions For CNN that evolved from traditional neural networks,
different activation functions will lead to different expression abilities of CNN,
especially between multi-layer connections. At present, nonlinear functions are
generally used as activation functions due to that nonlinear traits can be used to
improve the network’s limited approximation ability effectively. Common activation functions are sigmoid, tanh, rectified linear unit (ReLU), and the improved
activation functions based on ReLU, such as Leaky-ReLU, P-ReLU, and ReLU6.
The comparison between different activation functions is shown in Table 1.
Table 1. Comparison between different activation functions
Activation function Characteristics
sigmoid
(1) Gradient disappearance problem
(2) The output is not zero-centered
(3) Sigmoid contains power operation in the analytic
formula, which is time-consuming to solve
tanh
(1) The output is zero-centered
(2) Gradient disappearance problem still exist
ReLU
(1) Gradient disappearance problem is solved in the
positive range
(2) ReLU converges much faster than sigmoid and
tanh
(3) The output is not zero-centered
(4) Some neurons may never be activated, resulting
in the corresponding parameters that can never be
updated
Optimization Algorithms The optimization algorithms for deep learning
mainly include GD, SGD, Momentum, RMSProp, and Adam algorithms. Among
them, Momentum is an optimization algorithm based on the gradient-based
moving index weighted average. RMSProp calculates the differential squared
Précédent

- 245/679

Suivant