2.3 Stochastic Gradient Descent Method
25
The mystery of generalization in deep learning
On the other hand, as explained below, models in deep learning become more and
more complex with the number of layers. For example, the famous ResNet [24]
has hundreds of layers, which gives an incredibly huge d V C . 14 Therefore, from the
inequality of (2.11), it cannot be guaranteed that the generalization performance
is improved by the above logic. Nevertheless, ResNet appears to have gained
strong generalization performance. This appears to contradict with the above logic,
but it does not, because although the evaluation using (2.11) cannot guarantee
generalization performance, more detailed inequalities may exist:
(Generalization error)
?
≤ (Experience error) + O DL
something like this inside (2.11)?
≤ (Experience error) + O
log(#/d V C )
#/d V C
.
(2.17)
There is a movement to solve the mystery of the generalization performance of deep
learning from such more detailed inequality evaluation (for example, [27] etc.), but
there is no definitive result at the time of writing (as of January 2019).
2.3 Stochastic Gradient Descent Method
One of the easiest ways to implement learning is to use the gradient of the error
parameter J (the second term of the following equation) and update the parameters
as
J t +1 = J t − J D KL (P ||Q J ) ,
(2.18)
Where L is the maximum likelihood and k is the number of model parameters. Also, the AIC and
the amount of χ 2 that determine the accuracy of fitting have a relationship [23].
14 For example, according to the theorem 20.6 of [25], if the number of learning parameters of
a neural network having a simple step function as an activation function (described in the next
section) is N J , the VC dimension of the neural network is the order of N J log N J . ResNet is not
such a neural network, but let us estimate its VC dimension with this formula for reference. For
example, according to Table 6 of [24], a ResNet with 110 layers (=1700 parameters) has an average
error rate of 6.61% for the classification of CIFAR-10 (60,000 data), while the VC dimension is
12,645.25 according to the above formula, and the second term of (2.11) is 0.57. Since the errors
are scaled to [0,1] in the inequalities, the error rate can be read as at most about 10%, and the
above-mentioned error rate of 6.61% overwhelms this. A classification error of ImageNet [26]
(with approximately 10 7 data) is written in table 4 of the same paper, and this has a top-5 error
rate of 4.49% in 152 layers, while the same simple calculation gives about 9%, so the reality is
still better than the upper limit of the inequality. The inequality (2.11) is a formula for binary
classification, but CIFAR-10 has 10 classes and ImageNet has 1000 classes, so we should consider
the estimation here as a reference only.
25
The mystery of generalization in deep learning
On the other hand, as explained below, models in deep learning become more and
more complex with the number of layers. For example, the famous ResNet [24]
has hundreds of layers, which gives an incredibly huge d V C . 14 Therefore, from the
inequality of (2.11), it cannot be guaranteed that the generalization performance
is improved by the above logic. Nevertheless, ResNet appears to have gained
strong generalization performance. This appears to contradict with the above logic,
but it does not, because although the evaluation using (2.11) cannot guarantee
generalization performance, more detailed inequalities may exist:
(Generalization error)
?
≤ (Experience error) + O DL
something like this inside (2.11)?
≤ (Experience error) + O
log(#/d V C )
#/d V C
.
(2.17)
There is a movement to solve the mystery of the generalization performance of deep
learning from such more detailed inequality evaluation (for example, [27] etc.), but
there is no definitive result at the time of writing (as of January 2019).
2.3 Stochastic Gradient Descent Method
One of the easiest ways to implement learning is to use the gradient of the error
parameter J (the second term of the following equation) and update the parameters
as
J t +1 = J t − J D KL (P ||Q J ) ,
(2.18)
Where L is the maximum likelihood and k is the number of model parameters. Also, the AIC and
the amount of χ 2 that determine the accuracy of fitting have a relationship [23].
14 For example, according to the theorem 20.6 of [25], if the number of learning parameters of
a neural network having a simple step function as an activation function (described in the next
section) is N J , the VC dimension of the neural network is the order of N J log N J . ResNet is not
such a neural network, but let us estimate its VC dimension with this formula for reference. For
example, according to Table 6 of [24], a ResNet with 110 layers (=1700 parameters) has an average
error rate of 6.61% for the classification of CIFAR-10 (60,000 data), while the VC dimension is
12,645.25 according to the above formula, and the second term of (2.11) is 0.57. Since the errors
are scaled to [0,1] in the inequalities, the error rate can be read as at most about 10%, and the
above-mentioned error rate of 6.61% overwhelms this. A classification error of ImageNet [26]
(with approximately 10 7 data) is written in table 4 of the same paper, and this has a top-5 error
rate of 4.49% in 152 layers, while the same simple calculation gives about 9%, so the reality is
still better than the upper limit of the inequality. The inequality (2.11) is a formula for binary
classification, but CIFAR-10 has 10 classes and ImageNet has 1000 classes, so we should consider
the estimation here as a reference only.
