156
N. S. Philip
nodes from outside and everything about them are hidden. The hidden nodes also
have a transfer function (also known as activation function and is shown as green
in Fig. 3) that modulates the output before passing it to the subsequent layer. The
transfer function is a differentiable nonlinear function that also re-normalises the
output from each hidden node to a value between 0 and 1 or −1 and +1 depending
on the design of the model. The major purpose of the nonlinear transfer function
is to introduce nonlinearity in the network and to prevent the output of the node to
explode to large values due to the multiplication with the connection weights. The
backpropagation algorithm updates the connection weights by a small amount by
considering the derivative of the error with respect to the input. It then computes the
slope at each connection by applying chain rule on the partial derivatives with respect
to the connection weights. The slope is multiplied by a small quantity called learning
rate which increments the connection weight. The details will not be discussed here
as it is out of the scope of this article. Interested readers may refer [3] for a tutorial.
1.1 Adaptive Transfer Function
The transfer function, also called the activation function, plays a key role in the nonlinear behaviour of the ANN. Besides allowing to re-normalise the output of the node
to prevent factorial explosion of the output values, they also work as building blocks
to construct the high-dimensional nonlinear decision hyper surface that separates the
different classes of entities in the feature space. For the same reason, based on the
nature of the data, different transfer functions are in use to improve the predictive
accuracy of the ANN. Figure 4 gives a table of various activation functions that are
used in machine learning. As stated before, the primary criteria for the selection of
the activation function is that it should re-normalise the output from the node and
should be differentiable. Differentiable because backpropagation algorithm requires
derivatives to update the weights.
In 2002, Ninan and Joseph [4] proposed an adaptive transfer function that updates
its structure during the training process to optimise both learning speed and prediction
accuracy of the ANN. For this, the tanh(x) function was generalised to read as
AB F =
a + tanh(x)
1 + a
with a as a free adjustable parameter. The modification was shown to improve the
learning speed and accuracy of the trained network. In a later study [5], they demonstrated that the rainfall pattern in Trivandrum over the 87 years prior to 2003 could be
reliably predicted using ANN using adaptive transfer functions. This work got lot of
attention because it demonstrated the predictability of monsoon rains in Trivandrum
district prior to the dates when the chaotic changes due to global warming was not
visible.
Précédent

- 161/187

Suivant