52
3 Basics of Neural Networks
with dim[b 0 ] = N unit . Here assume that the activation function is not an even
function. A neural network with two hidden layers is
f 2h-NN (x) = j L · σ act (J σ act (j 0 x + b 0 ) + b 1 ) + b L .
(3.79)
Let us compare these and look at the meaning of deepening.
The case with a linear activation function
As a preparation, let us consider the case where the activation function is linear,
especially an identity map. That is,
σ act (x) = x .
(3.80)
In this case, for a single hidden layer, we find a linear map,
f 1h-NN (x) = j L · (j 0 x + b 0 ) + b L
(3.81)
= (j L · j 0 )x + (j L · b 0 + b L ) .
(3.82)
Furthermore, even with two layers, we find
f 2h-NN (x) = j L · (J 1 (j 0 x + b 0 ) + b 1 ) + b L
(3.83)
= (j L · J 1 j 0 )x + (j L · J 1 b 0 + j L · b 1 + b L ) .
(3.84)
In other words, if we make the activation function an identity map, we will always
get only a linear map.
The case with a nonlinear activation function
Next, consider the case where the activation function is nonlinear. In particular, we
want to consider the power as the complexity of a function that can be expressed by
a neural network, so we take
σ act (x) = x
3 .
(3.85)
For the sake of simplicity, assume that the neural network has a single hidden layer
with two units, j 0 = (j
0
1 , j
0
2 ) , j L = (j L
1 , j L
2 ) . Then we find
f 1h-NN (x) = j L · σ act ((j
0
1 x, j
0
2 x)
+ (b
0
1 , b
0
2 )
) + b L
(3.86)
= j L · σ act ((j
0
1 x + b
0
1 , j
0
2 x + b
0
2 )
) + b L
(3.87)
= j L ((j
0
1 x + b
0
1 )
3 , (j
0
2 x + b
0
2 )
3 )
+ b L
(3.88)
= j
L
1 (j
0
1 x + b
0
1 )
3
+ j
L
2 (j
0
2 x + b
0
2 )
3
+ b L .
(3.89)
3 Basics of Neural Networks
with dim[b 0 ] = N unit . Here assume that the activation function is not an even
function. A neural network with two hidden layers is
f 2h-NN (x) = j L · σ act (J σ act (j 0 x + b 0 ) + b 1 ) + b L .
(3.79)
Let us compare these and look at the meaning of deepening.
The case with a linear activation function
As a preparation, let us consider the case where the activation function is linear,
especially an identity map. That is,
σ act (x) = x .
(3.80)
In this case, for a single hidden layer, we find a linear map,
f 1h-NN (x) = j L · (j 0 x + b 0 ) + b L
(3.81)
= (j L · j 0 )x + (j L · b 0 + b L ) .
(3.82)
Furthermore, even with two layers, we find
f 2h-NN (x) = j L · (J 1 (j 0 x + b 0 ) + b 1 ) + b L
(3.83)
= (j L · J 1 j 0 )x + (j L · J 1 b 0 + j L · b 1 + b L ) .
(3.84)
In other words, if we make the activation function an identity map, we will always
get only a linear map.
The case with a nonlinear activation function
Next, consider the case where the activation function is nonlinear. In particular, we
want to consider the power as the complexity of a function that can be expressed by
a neural network, so we take
σ act (x) = x
3 .
(3.85)
For the sake of simplicity, assume that the neural network has a single hidden layer
with two units, j 0 = (j
0
1 , j
0
2 ) , j L = (j L
1 , j L
2 ) . Then we find
f 1h-NN (x) = j L · σ act ((j
0
1 x, j
0
2 x)
+ (b
0
1 , b
0
2 )
) + b L
(3.86)
= j L · σ act ((j
0
1 x + b
0
1 , j
0
2 x + b
0
2 )
) + b L
(3.87)
= j L ((j
0
1 x + b
0
1 )
3 , (j
0
2 x + b
0
2 )
3 )
+ b L
(3.88)
= j
L
1 (j
0
1 x + b
0
1 )
3
+ j
L
2 (j
0
2 x + b
0
2 )
3
+ b L .
(3.89)
