154
N. S. Philip
Fig. 1 The Biological Neuron takes input from external stimulus through its dendrites and the processed output is taken from the neuron to other neurons through the axon and chemical transporters
between the axon endings and dendrites of other neurons
scenario described by the input. This process of updating the connection weights is
what is called training. Like the biological neuron, a trained ANN will respond with
the same output when similar inputs are applied to it.
Although the possibility was well understood and was experimented in different
ways, it turned out to be impossible to adjust the weights except in trivial problems
until the backpropagation algorithm was developed [2]. Backpropagation algorithm
minimises the overall prediction error of the ANN by moving the weights along the
negative gradient of the error surface. To understand what this means, assume that
there is some value of the weight that correctly predicts the output. If the weight
is increased or decreased, there will be an error which will be negative on one
direction and positive on the other. Squaring the error makes it positive on either side
from the optimal value giving the error surface a parabolic shape. If random initial
weights are assigned and predictions are made, some will be predicted correctly
while some others will fail. In the case of all the failed predictions, their weights will
be somewhere on the parabola on either side of the optimal value for that weight.
[See Fig. 2] The gradient at the optimal location marked as the global cost minimum
location will be zero (because derivative at the point with respect to change in w
will be zero). Also the gradient is positive on the RHS (rising) of the minimum
and it is negative at the other side (falling). Thus, moving along the direction of the
negative gradient will always move in the direction of the minimum. That means,
gradually increasing or decreasing the weights (W) proportional to the negative of the
gradient will help to reach the minimum. Since we are climbing down the gradient in
both cases, it is called Gradient Descent Algorithm. The possible architecture of an
ANN is shown in Fig. 3. Each of the pink balls represent an input node of the ANN.
Traditionally, each such input is called a feature. For example, in a problem where
the goal is to differentiate apples from oranges, colour, shape, texture, etc., can be
Précédent

- 159/187

Suivant