classification accuracy and training time. Foody and Mathur (2006) also found that
SVM can produce a more accurate classification of cultivated landscape types.
Dixon and Candade (2008) compared SVM, MLC, and backpropagation neural
networks (NN) for classifying a Landsat scene, and found that SVM and NN
performed identically in the classification accuracy but SVM was more efficient
in the training phase. They also noted that SVM can be quite attractive when
working with high-dimensional data. This seems to be in line with an earlier
work conducted by Huang et al. (2002) who found that SVM performed better for
an image with seven bands than with three bands. The effectiveness of SVM for
working with high-dimensional data classification was also confirmed by several
other studies (e.g., Bazi and Melgani 2006; Camps-Valls et al. 2007), indicating
that they could provide a solution to dealing with the problem of “curse-ofdimensionality” (Hughes 1968). Although SVM have demonstrated strengths
when comparing with other classifiers, their performance can vary across different
land cover types (Foody and Mathur 2004a, b; Keramitsoglou et al. 2006; Su and
Huang 2009).
The performance of SVM can be affected by both parametric and
non-parametric factors (Foody and Mathur 2006; Yang 2011). Existing studies on
SVM classification have largely concentrated on either improving classification
accuracy on specific land cover types or reducing computational burdens, both of
which can be manipulated at the SVM configuration stage and at the training stage.
The inner-product kernel between the support vectors in feature space and in input
space largely determines the separability of optimal separable hyperplane (Haykin
1999). While introducing non-linear kernel functions could help deal with complex,
non-linear classification, it can also lead to the difficulty in choosing the most
appropriate kernel type and in the subsequent kernel parameterization (Huang
et al. 2002; Kavzoglu and Colkesen 2009; Yang 2011). Yang (2011) conducted
an empirical study assessing the performance of several most commonly used
kernel types, along with their internal parameterization, and found that the kernel
type and error penalty can substantially affect image classification accuracy. Some
customized kernels, particularly those incorporating both spatial and spectral information, were found to be quite promising when comparing with spectral-based
kernel types (Camps-Valls et al. 2006, 2007; Plaza et al. 2009).
Since the SVM is a supervised classifier by nature, both the size and quality of
training sample can affect the classification accuracy (Foody and Mathur 2006). For
land cover mapping from remote sensor imagery, training samples should consist of
relatively pure pixels, and should be identified from homogeneous areas in large
fields, which can be applicable for a variety of classifiers (Foody and Arora 1997).
SVM performance can be sensitive to the noise in training samples due to the use of
support vectors at the edges of class domains in feature space (Rodriguez-Galiano
et al. 2012). A minimum of 10–30 pixels per class per waveband should be used to
meet the assumption of normal distribution and be representative of the subclass
(Foody and Mathur 2004a, b, 2006). Like other non-parametric classifiers, there is
no need to maintain normal distributions in training samples for a SVM classification. Since only the support vectors are actually needed in constructing separate
268
D. Shi and X. Yang
SVM can produce a more accurate classification of cultivated landscape types.
Dixon and Candade (2008) compared SVM, MLC, and backpropagation neural
networks (NN) for classifying a Landsat scene, and found that SVM and NN
performed identically in the classification accuracy but SVM was more efficient
in the training phase. They also noted that SVM can be quite attractive when
working with high-dimensional data. This seems to be in line with an earlier
work conducted by Huang et al. (2002) who found that SVM performed better for
an image with seven bands than with three bands. The effectiveness of SVM for
working with high-dimensional data classification was also confirmed by several
other studies (e.g., Bazi and Melgani 2006; Camps-Valls et al. 2007), indicating
that they could provide a solution to dealing with the problem of “curse-ofdimensionality” (Hughes 1968). Although SVM have demonstrated strengths
when comparing with other classifiers, their performance can vary across different
land cover types (Foody and Mathur 2004a, b; Keramitsoglou et al. 2006; Su and
Huang 2009).
The performance of SVM can be affected by both parametric and
non-parametric factors (Foody and Mathur 2006; Yang 2011). Existing studies on
SVM classification have largely concentrated on either improving classification
accuracy on specific land cover types or reducing computational burdens, both of
which can be manipulated at the SVM configuration stage and at the training stage.
The inner-product kernel between the support vectors in feature space and in input
space largely determines the separability of optimal separable hyperplane (Haykin
1999). While introducing non-linear kernel functions could help deal with complex,
non-linear classification, it can also lead to the difficulty in choosing the most
appropriate kernel type and in the subsequent kernel parameterization (Huang
et al. 2002; Kavzoglu and Colkesen 2009; Yang 2011). Yang (2011) conducted
an empirical study assessing the performance of several most commonly used
kernel types, along with their internal parameterization, and found that the kernel
type and error penalty can substantially affect image classification accuracy. Some
customized kernels, particularly those incorporating both spatial and spectral information, were found to be quite promising when comparing with spectral-based
kernel types (Camps-Valls et al. 2006, 2007; Plaza et al. 2009).
Since the SVM is a supervised classifier by nature, both the size and quality of
training sample can affect the classification accuracy (Foody and Mathur 2006). For
land cover mapping from remote sensor imagery, training samples should consist of
relatively pure pixels, and should be identified from homogeneous areas in large
fields, which can be applicable for a variety of classifiers (Foody and Arora 1997).
SVM performance can be sensitive to the noise in training samples due to the use of
support vectors at the edges of class domains in feature space (Rodriguez-Galiano
et al. 2012). A minimum of 10–30 pixels per class per waveband should be used to
meet the assumption of normal distribution and be representative of the subclass
(Foody and Mathur 2004a, b, 2006). Like other non-parametric classifiers, there is
no need to maintain normal distributions in training samples for a SVM classification. Since only the support vectors are actually needed in constructing separate
268
D. Shi and X. Yang
