50
3 Basics of Neural Networks
3.3 Universal Approximation Theorem of Neural Network
Why can a neural network express the connection (correlation) between data? The
answer, in a certain limit, is the universal approximation theorem of neural
networks [32]. Similar theorems have been proved by many people, and here we
will explain a simple proof by M. Nielsen [33]. 9
Universal approximation theorem of one-dimensional neural network
Let us take the simplest model first. Consider a neural network with a single hidden
layer, that gives a one-dimensional real number when given a one-dimensional
variable x. It looks like this:
f (x) = J
(2)
· σ step (J
(1) x + b
(1) ) + b
(2) .
(3.72)
σ step (x) is step function which serves as an activation function,
σ step (x) =
0 (x < 0) ,
1 (x ≥ 0) .
(3.73)
When the activation function is a sigmoid function, it can be obtained at some limit
of it. 10 When the argument is multi-component, we define that the operation is on
each component. J (i) is the weight , and b (i) is the bias.
J
(l)
= (j
(l)
1 , j
(l)
2 , j
(l)
3 , · · · , j
(l)
n unit
)
,
(3.74)
b
(l)
= (b
(l)
1 , b
(l)
2 , b
(l)
3 , · · · , b
(l)
n unit
)
.
(3.75)
Here j
(l)
i and b (l) are real numbers. l is the serial number of the layer, l = 1, 2. n unit
is the number of units in each layer, which will be given later.
What happens at the first and second layers
First, let us see what happens at the first and second layers with n unit = 1. In this
case, the following is input to the third layer:
g 1 (x) = σ step (j
(1)
1 x + b
(1)
1 ) .
(3.76)
If we vary j
(1)
1 and b
(1)
1 , we find that the step function moves left or right. Then
multiply this by the coefficient j
(2)
1 ,
g 2 (x) = j
(2)
1 σ step (j
(1)
1 x + b
(1)
1 ) .
(3.77)
9 Visit his website for an intuitive understanding of the proof.
10 The situation is the same as the physics calculations at zero temperature at which the Fermi
distribution function becomes a step function.
Précédent

- 60/211

Suivant