CHAPTERS
Support Vector Machines
Mahesh Pal, Pakorn Watanachaturaporn
5.1
Introduction
Support Vector Machines (SVMs) are a relatively new generation of techniques
for classification and regression problems. These are based on Statistical Learning Theory having its origins in Machine Learning, which is defined by Kohavi
and Foster (1998) as,
... Machine Learning is the field of scientific study that concentrates
on induction algorithms and on other algorithms that can be said
to "learn."
The process of learning involves identification of a model based on the
training data that are used to make predictions about unknown data sets. Two
possible learning approaches can be employed (see Sect. 2.6 of Chap. 2): supervised and unsupervised. Both the learning approaches may be grouped into
two categories: parametric and non-parametric learning algorithms. Parametric learning algorithms such as the supervised maximum likelihood classifier
(MLC) require that the data be based on some pre-defined statistical model.
For example, the MLC assumes that the data follow Gaussian distribution. The
performance of a parametric learning algorithm, thus, depends on how well
the data match the pre-defined model and on the accurate estimation of model
parameters. Moreover, parametric learning algorithms generally suffer from
the problem called the curse of dimensionality, also known as the Hughes Phenomenon (Hughes 1968), especially under the situation when data dimensions
are very high such as the hyperspectral data from remote sensing sensors. The
Hughes phenomenon states that the ratio of the number of pixels with known
class identity (i. e. the training pixels) and the number of bands must be maintained at or above some minimum value to achieve statistical confidence. In
hyperspectral datasets, we have hundreds of bands, and it is difficult to have
a sufficient number of training pixels. Also, parametric learning algorithms
may have difficulty in classifying data at different measurement scales and
units. To overcome the limitations of parametric learning algorithms, several
non-parametric algorithms are in vogue for classification of remote sensing
data. These include nearest neighbor (Friedman 1994; Hastie and Tibshirani
P. K. Varshney et al., Advanced Image Processing Techniques for Remotely Sensed Hyperspectral Data
© Springer-Verlag Berlin Heidelberg 2004
Précédent

- 142/327

Suivant