13.2 Support Vector Machines
13.2.1 Basics
The basic idea behind the support vector machines (SVM) is to construct separating
hyperplanes between classes in feature space through the use of support vectors
which are lying at the edges of class domains; SVM seek the optimal hyperplane
that can separate classes from each other with the maximum margin (Vapnik 1995).
SVM were originally designed as a binary linear classifier, which assumes two
linearly separable classes to be partitioned. In most cases, the best separable
hyperplane may not be located exactly between two classes. To account for this,
an error item is introduced to manipulate the tradeoff between maximizing the
separation margin and minimizing the count of training samples that locates on the
wrong side. SVM are further extended to deal with non-linear classification by
using a non-linear kernel function to replace the inner product of optimal hyperplane. Several commonly used kernel functions include linear kernel, polynomial
kernel, radial basis function (RBF), and sigmoid kernel (Haykin 1999). Each of
these kernel functions is constructed with multiple parameters, and the parameter
settings can influence the performance of a specific support vector machine (Yang
2011).
Moreover, SVM have been used for multi-class mapping through reducing the
multi-class problem into a set of binary problems so that the basic SVM principles
can be still applied. Two commonly used strategies for this purpose include oneagainst-one and one-against-all (Foody and Mathur 2004b; Kavzoglu and Colkesen
2009). The former is generally preferred because of its less computational intensity
and comparable accuracy to the later. The one-against-all method can result in
unclassified instances (Huang et al. 2002; Hsu and Lin 2002; Pal and Mather 2005;
Mountrakis et al. 2011), which is not suitable for land cover mapping.
13.2.2 SVM for Land Cover Classification
The performance of SVM has been examined through some comparative studies
with other pattern classifiers for various land cover types (e.g., Huang et al. 2002;
Foody and Mathur 2006; Keramitsoglou et al. 2006; Su and Huang 2009). Huang
et al. (2002) found that SVM substantially outperformed maximum likelihood
(MLC) or decision tree (DC) in terms of classification accuracy and even surpassed
multilayer perceptron neural networks (MLP). Su and Huang (2009) implemented
SVM and MLC on a Multi-angle Imaging SpectroRadiometer (MISR) image to
differentiate eight semi-arid vegetation types, and found that SVM significantly
outperformed MLC. Keramitsoglou et al. (2006) mapped various vegetation types
using IKONOS data, and compared the performance of SVM with radial basis
(RBF) neural networks. They found that SVM had strengths in terms of
13 Support Vector Machines for Land Cover Mapping from Remote Sensor Imagery
267
13.2.1 Basics
The basic idea behind the support vector machines (SVM) is to construct separating
hyperplanes between classes in feature space through the use of support vectors
which are lying at the edges of class domains; SVM seek the optimal hyperplane
that can separate classes from each other with the maximum margin (Vapnik 1995).
SVM were originally designed as a binary linear classifier, which assumes two
linearly separable classes to be partitioned. In most cases, the best separable
hyperplane may not be located exactly between two classes. To account for this,
an error item is introduced to manipulate the tradeoff between maximizing the
separation margin and minimizing the count of training samples that locates on the
wrong side. SVM are further extended to deal with non-linear classification by
using a non-linear kernel function to replace the inner product of optimal hyperplane. Several commonly used kernel functions include linear kernel, polynomial
kernel, radial basis function (RBF), and sigmoid kernel (Haykin 1999). Each of
these kernel functions is constructed with multiple parameters, and the parameter
settings can influence the performance of a specific support vector machine (Yang
2011).
Moreover, SVM have been used for multi-class mapping through reducing the
multi-class problem into a set of binary problems so that the basic SVM principles
can be still applied. Two commonly used strategies for this purpose include oneagainst-one and one-against-all (Foody and Mathur 2004b; Kavzoglu and Colkesen
2009). The former is generally preferred because of its less computational intensity
and comparable accuracy to the later. The one-against-all method can result in
unclassified instances (Huang et al. 2002; Hsu and Lin 2002; Pal and Mather 2005;
Mountrakis et al. 2011), which is not suitable for land cover mapping.
13.2.2 SVM for Land Cover Classification
The performance of SVM has been examined through some comparative studies
with other pattern classifiers for various land cover types (e.g., Huang et al. 2002;
Foody and Mathur 2006; Keramitsoglou et al. 2006; Su and Huang 2009). Huang
et al. (2002) found that SVM substantially outperformed maximum likelihood
(MLC) or decision tree (DC) in terms of classification accuracy and even surpassed
multilayer perceptron neural networks (MLP). Su and Huang (2009) implemented
SVM and MLC on a Multi-angle Imaging SpectroRadiometer (MISR) image to
differentiate eight semi-arid vegetation types, and found that SVM significantly
outperformed MLC. Keramitsoglou et al. (2006) mapped various vegetation types
using IKONOS data, and compared the performance of SVM with radial basis
(RBF) neural networks. They found that SVM had strengths in terms of
13 Support Vector Machines for Land Cover Mapping from Remote Sensor Imagery
267
