142
Biomedical Signal and Image Processing
to a particular class and −1 if the sample does not belong to this class. A perceptron
first initializes the weights of the network by random numbers and then uses an
iterative technique for updating of the weights, i.e., when encountering an example,
the values of the weights are iteratively adjusted to produce the correct or better classification. Specifically, assume that for a sample x j = (x j1 , x j2 ,…, x jn ) the true target
output is t. We use this knowledge to update the weights. Assume that with the old
(current) values of the weights, the estimated output for x j is y. Then, if y = t, then the
old values of weights are fine and no real update is needed. Otherwise, the old values
of the weights are updated as follows:
w i (new ) = w i (old ) + tx j i
(7.17)
b n
( ew ) b ( old ) t
=
+
where
x ji is the ith element of the sample vector x j
w i (new) is the new value of the weight w i
The training process is continued until the exposure of the network can correctly
classify all examples in the training set. In order to train a neural network of any
type, it is often the case that the network must be exposed to all training quite a
few times. Each cycle of exposing the network to all training examples is called
an epoch.
The learning process described in the following can result to rather oscillatory
results in consecutive epochs. In other words, each time an example is used for training, the weights are significantly changed to accommodate that particular example,
but as soon as the next example is used for training, the weights are very much
changed in favor of the new example. The same scenario is repeated in the consecutive epochs, and, therefore, the values of weights oscillate among some sets
that accommodate particular examples. In order to address this issue, instead of the
learning equations of Equation 7.15, the following learning criteria are used:
w i (new ) = w i (old ) + atx j i
(7.18)
b n
( ew ) = b ( old )
t
+ a
In the preceding equations, 0 < α < 1 is called the learning rate. The fact that the
learning rate is less than one, each time a new example is used for training, instead
of adjusting the weights completely to accommodate that training example as we did
in the previous training scheme, the weights are adjusted only to some degree that
depends on the exact value of α. Choosing α to be a number close to 0.1 or so would
effectively eliminate the unwanted oscillations in the weight values.
Before providing an example for perceptron, we focus on the perceptron’s resulting classifier. A closer look at the formulation of perceptron reveals that a perceptron
Précédent

- 169/412

Suivant