2.3 Stochastic Gradient Descent Method
27
Calculate ˆ
g according to the formula (2.22)
J t +1 = J t − ˆ
g.
(2.23)
Note that in order to justify this approximation using the law of large numbers, all
sampling in (2.21) must be taken independently, so it must be performed separately
at each step t of the parameter update. Such a learning method that always uses new
data at each t is called online learning.
When data is limited
When a prepared database is used, the above is not possible. In that case, a method
called batch learning is used. We want to consider a stochastic gradient descent
method that also minimizes the bias, and in deep learning, the following method
called mini-batch learning is often used:
(Data{(x[i], d[i])} i=1,2,...# shall be given in advance)
1. Initialize J properly.
2. Repeat the following:
Divide data intoM partial data randomly.
Repeat the following from m = 1 to m = M:
Replace data in (2.22) with the mth partial data and calculate ˆ
g.
J ← J − ˆ
g.
(2.24)
Note that at the beginning of the loop in step 2, there is a process for randomly
splitting the data. This method is mostly used when applying the gradient descent
method in supervised learning. Unlike the original gradient descent method (2.23),
even if the loop of step 2 is repeated, each ˆ
g is not independent of the others,
thus there is no guarantee that the approximation accuracy of the gradient of the
generalization error will be better. Normally, it is necessary to prepare validation
data separately and monitor (observe) the empirical error at the same time, so as to
avoid over-training.
The case of deep learning
In most cases, the training is done according to the method (2.24). Here are some
points to emphasize. First, the gradient ˆ
g is the gradient of the error function
which will be introduced later, and when a neural network is used, the differential
calculation can be algorithmized: it is called back-propagation, and provides a fast
algorithm which will be explained later. Also, be aware that not only the simple
stochastic gradient descent method but also various evolved forms of the gradient
descent method are often used. References that summarize various gradient methods
include [28]. In addition, original papers (listed in [28]) can be read for free.
Précédent

- 37/211

Suivant