134
5: Mahesh Pal, Pakorn Watanachaturaporn
1996; Mather 1999), decision tree (Richards and Jia 1999; pal 2002) and neural network (Haykin 1999; Tso and Mather 2001) algorithms. The latter two
fall in the category of machine learning algorithms. Neural network classifiers
are sometimes touted as substitutes for the conventional MLC algorithm in
the remote sensing community. The preferred neural network classifier is the
feed-forward multi-layer perceptron learnt with a back-propagation algorithm
(see Sect. 2.6.3 of Chap. 2). However, even though neural networks have been
successful in classifying complex data sets, they are slow during the training
phase. A number of studies have also reported that neural network classifiers have problems in setting various parameters during training. Moreover,
these may also have limitations in classifying hyperspectral datasets since the
complexity of the network architecture increases manifolds. Nearest-neighbor
algorithms are sensitive to the presence of irrelevant parameters in the dataset
such as noise in a remote sensing image. In case of decision tree classifiers, as
the dimensionality of the data increases, class structure becomes dependent
on a combination of features thereby making it difficult for the classifier to
perform well (Pal 2002).
Recently, the Support Vector Machine (SVM), another machine learning
algorithm, has been proposed that may overcome the limitations of aforementioned non-parametric algorithms. SVMs, first introduced by Boser et al.
(1992) and discussed in more detail byVapnik (1995,1998), have their roots in
statistical learning theory (Vapnik 1999) whose goal is to create a mathematical framework for learning from input training samples with known identity
and predict the outcome of data points with unknown identity. This results in
two important theories. The first theory is called empirical risk minimization
(ERM) where the aim is to minimize the learning or training error. The second
theory is called structural risk minimization (SRM), which is aimed at minimizing the upper bound on the expected error over the whole dataset. SVMs
are based on the SRM theory while neural networks are based on ERM theory.
An SVM is basically a linear learning machine based on the principle of optimal separation of classes. The aim is to find a linear separating hyperplane that
separates classes of interest. The hyperplane is a plane in a multidimensional
space and is also called a decision surface or an optimal separating hyperplane
or an optimal margin hyperplane. The linear separating hyperplane is placed
between classes in such a way that it satisfies two conditions. First, all the data
vectors that belong to the same class are placed on the same side of the hyperplane. Second, the distance or margin between the closest data vectors in both
the classes is maximized (Vapnik and Chervonenkis 1974; Vapnik 1982). In
other words, the optimum hyperplane is the one that provides the maximum
margin between the two classes. For each class, the data vectors forming the
boundary of classes are located on supporting hyperplanes - the term used
in the theory of convex sets. Thus, these data vectors are called the Support
Vectors (Scholkopf 1997). It is noteworthy that the data vectors located along
the class boundary are the most significant ones for SVMs.
Many times, a linear separating hyperplane is not able to classify input data
without error. Under such circumstances, the data are transformed to a higher
dimensional space using a non-linear transformation that spreads the data
Précédent

- 143/327

Suivant