118
4: Stefan A. Robila
The second term in (4.25) can be further developed:
taE{log(P(ui))} = tE!alog(p(Ui))) = t E !_I_a p (Ui))
i=1
aw
i=1
aw
i=1
P (Ui) aw
= t E 1_1_ ap (Ui) aUi ).
(4.27)
i=1
p (Ui) aUi aw
The derivative of Ui with respect to the (i,j) element of w, (i. e. Wj,k) can be
computed as follows:
(
n
)
a L Wi,qXq
n
.
. .
q=1
= L a (Wi,qXq) = {Xk if 1 = J
aWl· k
aWl· k
0 if i =J j .
,
q=1
'
(4.28)
Now, define a family of nonlinear functions gi(·) such that they approximate
the probability density function for each Ui. Using the above result, the second
term in (4.25) becomes:
t aE {log (p (Ui))} ~ t E 1_1_ agi (Ui) aUj ) = E {_I ag(u) xT} ,
i=1
aw
i=1
gi (Ui) aUi aw
g(u) au
(4.29)
where g(u) denotes (g1 (ud, ... ,gn(un)). Finally, compute an approximation of
the gradient of 1(·) with respect to W and obtain the update step:
.1W ~ - aI(u) = (W-1)T + E {(_l_a g (U)) . xT } .
aw
g(u) au
( 4.30)
In the form given in (4.30), a matrix inversion is required. This may result
in slow processing. To speed up, the update step is multiplied by the 'natural
gradient' WTW (Lee 1998):
.1W ~ - aI(u) WTW = (W-1)T WTW + E {(_l_a g (U)) xTWT } W
aw
~~ ~
where
h(u) = _1_ ag(u) .
g(u) au
= W + E {h(u)u T } W,
(4.31)
It can be proved that multiplication with the natural gradient not only
preserves the direction of the gradient but also speeds up the convergence
Précédent

- 127/327

Suivant