Overview of Image Processing
73
also result in the phenomenon called overtraining (Chi 1995). The optimal
number of hidden units is often determined experimentally using a trial and
error method although some basic geometrical arguments may be used to give
an approximate indication (Bernard et al. 1995).
Thus, the units in the neural network are representation of the biological
concept of a neuron and weighted paths connecting them (Schalkoff 1997).
The data supplied to a unit's input are multiplied by the path's weight and
are summed to derive the net input to that unit. The net input (NET) is then
transformed by an activation function (f) to produce an output for the unit. The
most common form of the activation function is a Sigmoid function, defined
as
1
t(NET) = - - - -
1+ exp( -NET)
and accordingly,
Output = t(NET) ,
(2.25)
(2.26)
where NET is the sum of the weighted inputs to the processing unit and may
be expressed as
n
NET = LXiWi,
;=1
(2.27)
where Xi is the magnitude of the ith input and Wi is the weight of the interconnected path (Foody 1998).
The determination of the appropriate weights is referred to as learning
or training. Learning algorithms may be supervised and unsupervised, as
discussed earlier. Generally a backpropagation algorithm is applied which
is a supervised algorithm and has been widely used in applications related
to neural network classification in remote sensing. The algorithm iteratively
minimizes an error function over the network outputs and target outputs
for a sample of training pixels c. The process continues until the error value
converges to a minimum. Conventionally, the error function is given as
c
E = 0.5 L (Ti - Oi)2 ,
(2.28)
i=1
where Ti is the target output and Oi is the network output, also known as the
activation levels.
Thus, the magnitudes of the weights are determined by an iterative training
procedure in which network repeatedly tries to learn the correct output for
each training sample. The procedure involves modifying the weights of the
layers connecting units until the network is able to characterize the training
data (Foody 1995b). Once the network is trained, the adjusted weights are used
to classify the unknown dataset.
73
also result in the phenomenon called overtraining (Chi 1995). The optimal
number of hidden units is often determined experimentally using a trial and
error method although some basic geometrical arguments may be used to give
an approximate indication (Bernard et al. 1995).
Thus, the units in the neural network are representation of the biological
concept of a neuron and weighted paths connecting them (Schalkoff 1997).
The data supplied to a unit's input are multiplied by the path's weight and
are summed to derive the net input to that unit. The net input (NET) is then
transformed by an activation function (f) to produce an output for the unit. The
most common form of the activation function is a Sigmoid function, defined
as
1
t(NET) = - - - -
1+ exp( -NET)
and accordingly,
Output = t(NET) ,
(2.25)
(2.26)
where NET is the sum of the weighted inputs to the processing unit and may
be expressed as
n
NET = LXiWi,
;=1
(2.27)
where Xi is the magnitude of the ith input and Wi is the weight of the interconnected path (Foody 1998).
The determination of the appropriate weights is referred to as learning
or training. Learning algorithms may be supervised and unsupervised, as
discussed earlier. Generally a backpropagation algorithm is applied which
is a supervised algorithm and has been widely used in applications related
to neural network classification in remote sensing. The algorithm iteratively
minimizes an error function over the network outputs and target outputs
for a sample of training pixels c. The process continues until the error value
converges to a minimum. Conventionally, the error function is given as
c
E = 0.5 L (Ti - Oi)2 ,
(2.28)
i=1
where Ti is the target output and Oi is the network output, also known as the
activation levels.
Thus, the magnitudes of the weights are determined by an iterative training
procedure in which network repeatedly tries to learn the correct output for
each training sample. The procedure involves modifying the weights of the
layers connecting units until the network is able to characterize the training
data (Foody 1995b). Once the network is trained, the adjusted weights are used
to classify the unknown dataset.
