3.3 Universal Approximation Theorem of Neural Network
51
Then we can freely change the height (even in the negative direction). Adding a
bias term b
(2)
1 for the second layer can also change the minimum value of the step
function.
Then, let n unit = 2. We can draw a rectangle function of any shape using J (l) and
b (l) as parameters. Now we are ready for the proof. Consider n unit = 2k (k > 1).
Then we can draw a graph with many rectangle functions superimposed.
Given a continuous objective function t (x), the neural network f (x) considered
above can approximate it by increasing n unit and adjusting the parameters. One
can get as close as one likes. 11 Therefore, a one-dimensional neural network can
represent any continuous function.
Universal approximation theorem of neural networks
The facts mentioned above can be extended to higher dimensions. That is, a
continuous mapping from x to f(x) can be expressed using a neural network. This
fact is called the universal approximation theorem of neural networks.
At the same time, note that this theorem refers to an unrealistic situation. First, the
activation function is a step function. In the learning based on the backpropagation
method, the activation function is required to be differentiable, so the step function
cannot meet the purpose. Second, we need an infinite number of intermediate units.
As a quantitative problem, it does not describe how the neural network improves
its approximation to the target function at what speed and with what accuracy.
Another caveat is that this theorem only says that a neural network can give an
approximation of the desired function. Nevertheless, this theorem, to some extent,
intuitively explains a fragment of why neural networks are powerful.
Why deep layers?
So far, we have seen a neural network with a single hidden layer. Even if the number
of hidden layers is one, if the number of units is infinite, any function can be
approximated. Then, why is deep learning effective?
The following facts are known [34]:
1. More units in the middle layer ⇒ Expressivity increases in power
2. Deeper neural network ⇒ Expressivity increases exponentially
In this section, we will consider why deepening is useful, using a toy model.
Toy model of deep learning
Let us introduce a simple toy model that can show that the number of hidden layers
affects the expressivity of the neural network. A neural network with one hidden
layer is
f 1h-NN (x) = j L · σ act (j 0 x + b 0 ) + b L ,
(3.78)
11 The part corresponds to the dense nature in Cybenko’s proof.
51
Then we can freely change the height (even in the negative direction). Adding a
bias term b
(2)
1 for the second layer can also change the minimum value of the step
function.
Then, let n unit = 2. We can draw a rectangle function of any shape using J (l) and
b (l) as parameters. Now we are ready for the proof. Consider n unit = 2k (k > 1).
Then we can draw a graph with many rectangle functions superimposed.
Given a continuous objective function t (x), the neural network f (x) considered
above can approximate it by increasing n unit and adjusting the parameters. One
can get as close as one likes. 11 Therefore, a one-dimensional neural network can
represent any continuous function.
Universal approximation theorem of neural networks
The facts mentioned above can be extended to higher dimensions. That is, a
continuous mapping from x to f(x) can be expressed using a neural network. This
fact is called the universal approximation theorem of neural networks.
At the same time, note that this theorem refers to an unrealistic situation. First, the
activation function is a step function. In the learning based on the backpropagation
method, the activation function is required to be differentiable, so the step function
cannot meet the purpose. Second, we need an infinite number of intermediate units.
As a quantitative problem, it does not describe how the neural network improves
its approximation to the target function at what speed and with what accuracy.
Another caveat is that this theorem only says that a neural network can give an
approximation of the desired function. Nevertheless, this theorem, to some extent,
intuitively explains a fragment of why neural networks are powerful.
Why deep layers?
So far, we have seen a neural network with a single hidden layer. Even if the number
of hidden layers is one, if the number of units is infinite, any function can be
approximated. Then, why is deep learning effective?
The following facts are known [34]:
1. More units in the middle layer ⇒ Expressivity increases in power
2. Deeper neural network ⇒ Expressivity increases exponentially
In this section, we will consider why deepening is useful, using a toy model.
Toy model of deep learning
Let us introduce a simple toy model that can show that the number of hidden layers
affects the expressivity of the neural network. A neural network with one hidden
layer is
f 1h-NN (x) = j L · σ act (j 0 x + b 0 ) + b L ,
(3.78)
11 The part corresponds to the dense nature in Cybenko’s proof.
