acceptance domain of the previous layer, and then extracts features of that acceptance
domain [11]. The process of forward transmission is shown in Formula 1.
f out ¼ f w. . . wf wx þ b
ð
Þþb
ð
Þ . . . þ b
ð
Þ
ð 1Þ
The back-propagation process is actually a process of weight update. In this paper,
SGD and Adam are used to update the weight. The process of feature extraction and
weight update is continuous, until the global optimal solution is found. In most
experiments, the results are local optimum; it is necessary to adjust the learning rate and
select the loss function, so that the network can find the global optimum solution.
The parameter update is shown in Formula (2).
W
l
ij ¼ W
l
ij À a
@
@W
ðlÞ
ij
JðW; bÞ
ð 2Þ
5 Binarized Convolutional Neural Network
In the article [12], the parameters with single-precision floating point are converted into
parameters that only occupy 1 bit, which theoretically reduces the memory space by 32
times. Moreover, the speed of the model is accelerated to about twice as fast. Symbolic
functions can be expressed as:
signðxÞ ¼
1
ifx [ 0
À1 if x 0
&
ð3Þ
The binarized convolution neural network converts the weights and the activation
values of hidden layers into 1 or −1. There are two ways of conversion: deterministic
method and stochastic method. The former is simple and intuitive. If the weight or
activation value is greater than 0, it is converted to 1; if it is less than 0, it is converted
to −1. The latter calculates a probability p for the input. When p is greater than a
threshold, it is +1, otherwise it is −1. Since the random numbers generated by hardware, which is more difficult to implement. Therefore, the first method is adopted in
this paper.
In this paper, we take a method of approximating the real weights by using a
binarized weight B and a scale factor a. The process of conversion is shown in Formula (4), where ⊕ represents the convolution operation of input and binarization
weights.
I Ã W % I È B
ð
Þa
ð4Þ
The binarized VGG16 network is used for experiments in this paper. Same as
ordinary convolutional neural networks, the binarized network also includes the input
layer, the hidden layer, and the output layer. The hidden layer includes convolution
layer and pooling layer, where binarized convolution and pooling operation are used.
14
X. Pu et al.
domain [11]. The process of forward transmission is shown in Formula 1.
f out ¼ f w. . . wf wx þ b
ð
Þþb
ð
Þ . . . þ b
ð
Þ
ð 1Þ
The back-propagation process is actually a process of weight update. In this paper,
SGD and Adam are used to update the weight. The process of feature extraction and
weight update is continuous, until the global optimal solution is found. In most
experiments, the results are local optimum; it is necessary to adjust the learning rate and
select the loss function, so that the network can find the global optimum solution.
The parameter update is shown in Formula (2).
W
l
ij ¼ W
l
ij À a
@
@W
ðlÞ
ij
JðW; bÞ
ð 2Þ
5 Binarized Convolutional Neural Network
In the article [12], the parameters with single-precision floating point are converted into
parameters that only occupy 1 bit, which theoretically reduces the memory space by 32
times. Moreover, the speed of the model is accelerated to about twice as fast. Symbolic
functions can be expressed as:
signðxÞ ¼
1
ifx [ 0
À1 if x 0
&
ð3Þ
The binarized convolution neural network converts the weights and the activation
values of hidden layers into 1 or −1. There are two ways of conversion: deterministic
method and stochastic method. The former is simple and intuitive. If the weight or
activation value is greater than 0, it is converted to 1; if it is less than 0, it is converted
to −1. The latter calculates a probability p for the input. When p is greater than a
threshold, it is +1, otherwise it is −1. Since the random numbers generated by hardware, which is more difficult to implement. Therefore, the first method is adopted in
this paper.
In this paper, we take a method of approximating the real weights by using a
binarized weight B and a scale factor a. The process of conversion is shown in Formula (4), where ⊕ represents the convolution operation of input and binarization
weights.
I Ã W % I È B
ð
Þa
ð4Þ
The binarized VGG16 network is used for experiments in this paper. Same as
ordinary convolutional neural networks, the binarized network also includes the input
layer, the hidden layer, and the output layer. The hidden layer includes convolution
layer and pooling layer, where binarized convolution and pooling operation are used.
14
X. Pu et al.
