3.2 Derivation of Backpropagation Method Using Bracket Notation
49
= =δ 1 |δJ N |h N−1 + +δ 1 |J N
δ|h N−1
Transform this with (3.60)
= =δ 1 |δJ N |h N−1 + +δ 1 |J N G N−1
Define this =::δ 2 |
δJ N−1 |h N−2 + J N−1 δ|h N−2
= · · ·
= =δ 1 |δJ N |h N−1 + +δ 2 |δJ N−1 |h N−2 + · · · + +δ N |δJ 1 |h 0 .
(3.65)
Here the last |h 0 =
m |m =
m |m m is a ket representation of x.
Therefore, if we want to reduce the value of the error function, we just take
δJ l = − N−l+1 l−1 | .
(3.66)
This is because with this we find
δE = −
|δ 1
2
|h N−1
2 +
|δ 2
2
|h N−2
2 + . . .
|δ N
2
|h 0
2
,
(3.67)
and the change of E becomes a small negative number. We can repeat this with
various pairs (x, d). By the way, the right-hand side of (3.66) is minus times
the differentiation of the error function E, so this is a stochastic gradient descent
method described in the previous chapter. Bra vector l | artificially introduced in
the derivation of (3.65) satisfies the recurrence formula
l | = =δ l−1 |J N−l+2 G N−l+1 ,
(3.68)
with J N+1 = 1, which is an expression “similar” to (3.59). The backpropagation
algorithm is an algorithm that calculates the differential value of E for all J l by
combining (3.59), (3.66) and (3.68). It is
1. Repeat (3.59) with an initial value |h 0 = |x[i]], and calculate |h l
2. Repeat (3.68) with an initial value 0 | = =d[i]|, and calculate l |.
3. ∇ J l E = |δ N−l+1 l−1 |.
(3.69)
The name “backpropagation” means that the propagation direction of by (3.68)
is opposite to the direction of (3.59). This fact is written in our bracket notation as
|h : ket : Forward propagation,
(3.70)
: bra : Backpropagation.
(3.71)
The correspondence is visually clear.
Précédent

- 59/211

Suivant