148
Biomedical Signal and Image Processing
These errors in output units are then distributed back to the units of hidden layer
and are also used to update the weights between hidden layer units and output
layer units. In order to see how this is done, note that for the neurons in the hidden
layers we have
z 0j = v oj +
x i i
v j
(7.28)
∑
i
and
z j = f z 0 j
( )
(7.29)
where
v j s are the weights ending at the hidden unit j
i
v j is the bias weight of the hidden unit j
o
z 0j is the weighted sum of the inputs entering the hidden node j
z j is the output of the hidden node j
The updating of the weights v j s and v j is then performed as follows:
i
o
v t
( 1) v t
ij
+ ad x
(7.30)
ij
+ = ( )
j i
v t
( + =
1) v t
( ) + ad
(7.31)
0 j
0 j
j
In the preceding equations, α is the learning rate as defined for the perceptron. Note
the role δ j plays in the updating of v j s and v j . This role described why the algorithm
i
o
is named as backpropagation; the error of the output layer is for updating of the
weights of the hidden layer. The same process is used to update the weights between
input units and any other hidden layer units. The training process continues until
either the total error at the output is not improving or a prespecified maximum number
of iterations have been passed.
7.7.2.3 Momentum
While training a neural network, the training algorithm often encounters a noisy
sample. Algorithms such as backpropagation are often misled by the noisy data and
change the weights according to the noisy or even wrong data. The noisy data also
cause the training process to take much longer to reach the desired set of parameters.
This problem encourages us to, instead of updating weights only based on new sample,
consider the momentum of the previous values of weight in updating the weights. If the
previous examples (that are less noisy) are pushing weights in the right direction,
the momentum of the previous samples can somehow avoid the undesirable effects
of the noisy data on the weights.
In order to use momentum in backpropagation algorithm, we must consider
weights from one or more previous training patterns. For example, in the simplest
Précédent

- 175/412

Suivant