170
Activation functions, g(⋅), a set of non-linear transformations are required to suitably
extract features. They are applied element-wise over each feature map in the hidden layers.
Some activation functions are Sigmoid, TanH and ReLU. Each activation function has an
active zone, that is, an interval where the derivative of the function is not zero.
The Batch Normalization function (BN) (Ioffe & Szegedy, 2015) normalizes (
)
W H
b
m
m
W H
W
m
b
− +
before passing it through the activation function by subtracting the mean and dividing by their
respective standard deviation over each feature map. This process has three main properties:
(1) it avoids vanishing gradient problems during training, by adjusting the input values to the
active zone; (2) it accelerates the training process; and (3) it serves as a regularization method.
The pooling function, pool ( )
⋅ , seeks to reduce the size of the representative hidden layer
by taking small regions of each feature map. Several functions exist including the widely used
max pooling, and others such as average pooling, min pooling and L2-norm pooling. They are
usually applied to go from a moving window of (2, 2) to a single value (size (1, 1). Max pooling keeps only the maximum value among the nodes in the small region. This pooling function significantly reduces the number of learning parameters, improving statistical efficiency
and reducing the memory storage consumption (Goodfellow et al., 2016).
The classification section is composed by a set of feedforward networks (Hornik et al.,
1989) whose building blocks are:
W FC
W W
f
C : weight matrix. Dimensions (
)
n n
FC
FC
( )
f
( )
f −
corresponding to the previous fully connected layer size and the next fully connected desired size.
B f :
bias vector.
FC f : hidden fully connected layer. Dimensions ( )
n FC
( )
f
.
)
where FC f results of a matrix multiplication between W FC
W W
f
C
and FC f −1 , adding a bias vector
b f and then passing the temporary result only by (1) a Batch-Normalization function and
(2) an activation function as:
FC f
Vec
f
g
f F
=
( )
M
H M
=
(
)
BN (
)
FC
f
f
W FC
b
FC
W W
f
f
b
f
C C
≤
f
⎧
⎨
⎪
⎧ ⎧
⎨ ⎨
⎩ ⎪
⎨ ⎨
⎩ ⎩
if
i
0
(2)
The first feedforward network FC 0 corresponds to a vector-representation Vec ( )
M
H
of
the last hidden layer in the feature extraction block. The BN and activation functions act
exactly as presented before.
When categorical distributions are required, the final layer FC F must have the same
dimensions as the number of categories. Let K be the number of categories, n
K
FC
( )
F
, and
s k ∈{
}
s
s K
s
s the non-normalized conditional probability of each class in FC F . By passing
FC F through a softmax function:
( )
1
ˆ
k
c
s
K s
c
e
p k
(
ˆ
e
=
= ∑
(3)
the expected conditional probabilities for each class are obtained. When P continuous points
are required to be predicted, a usual technique is to apply another fully connected layer
FC F = g (
)
W
FC
b
C F
C
W
FC
(
)
1 = g
b F
b +
b
F
W W C
F
FC
F
b
C
W
FC
W W
FC
W W
FC
where W FC
W W
F
C +1
has dimensions ( )
y y
P n FC
,
( )
F
( ) and the activation function depends on the nature of the predicted variable, that is, if the expected output is in the
range
s
[ ]
0,s ∈
+
R a ReLU function seems suitable, while outputs with a normal gaussian
distribution would be better treated with a Sigmoid or TanH function.
Inner parameters (
)
Θ = { } are initialized as random values from a normal gaussian
truncated function. Optimum Θ values are obtained as result of training the CNN by applying
the Adam optimizer (Kingma & Ba, 2014) algorithm. The loss functions to minimize during
training (Eq. (4)) for categorical variables, considering the predicted probability of each class
(Eq. (3)) and the real probability p(k), is the negative log-likelihood of the Cross Entropy
Précédent

- 191/780

Suivant