Support Vector Machines for Classificationof Multi- and Hyperspectral Data
239
10.2
Parameters Affecting SVM Based Classification
Application of SVMs to any classification and regression problem requires
the determination of several design parameters and the resolution of several
issues. Some of the parameters and issues considered here are:
1. Determination of the penalty value (C value) to regulate the error term
(see Sect. 5.3.2)
2. Choice of an appropriate kernel and its parameters (see Sect. 5.3.3).
3. Selection of a suitable multiclass method (see Sect. 5.4).
4. Decision on the type of optimizer to be used (see Sect. 5.5).
In practice, the training data do not comprise of only pure pixels. Some of
them are mixed pixels due to either the noise present in the training data or
the mixture of classes during the selection of training data due to mislabeling
thereby introducing classification errors. A penalty parameter is introduced to
handle this error. The penalty parameter is also called the C value and depends
on the training data and the type of kernel. In general, it is found by trial and
error.
A kernel function determines the characteristic of an SVM. When the classes
in the input data are separable using a linear hyperplane, a linear kernel is used.
Similarly, a nonlinear kernel function may be used to separate data when the
classes are nonlinearly separable. However, it is very difficult to determine
which non-linear kernel function is suitable for a given application. In practice, a user will apply standard kernel functions such as a linear, polynomial,
RBF (also known as the Gaussian kernel), or sigmoid kernel. If these standard
kernel functions are unable to perform adequately, a user may have to design
a problem-specific kernel function. Moreover, the choice of the kernel further
brings in another issue that a user has to consider, which is to determine the values of parameters for a given kernel function. These parameters may be called
hyperparameters. For example, in case of an RBF kernel, the hyperparameter
that needs to be determined is a (see Sect. 5.3.3).
Thus, finding the parameters or hyperparameters that are able to accurately
predict the unknown data for a given application requires a method of model
selection or parameter search. The unknown data, in this case, refers to testing
samples or the image to be classified during the allocation stage. It is also a fact
that a model that achieves high training accuracy may not guarantee high
accuracy for the unknown data. Therefore, to establish that the model works
well on the whole dataset (i. e. the model has high generalization capacity),
a common strategy is to split the data into two parts - a training set, and
a validation set or a testing set. An SVM is trained on the training set and is
evaluated for its accuracy on the testing set. Another model selection method
is called k-fold cross-validation. In this method, the training data are divided
into k subsets of equal size. An SVM is trained on the k - 1 subsets of data
and is tested on the remaining one subset. Training and testing are performed
239
10.2
Parameters Affecting SVM Based Classification
Application of SVMs to any classification and regression problem requires
the determination of several design parameters and the resolution of several
issues. Some of the parameters and issues considered here are:
1. Determination of the penalty value (C value) to regulate the error term
(see Sect. 5.3.2)
2. Choice of an appropriate kernel and its parameters (see Sect. 5.3.3).
3. Selection of a suitable multiclass method (see Sect. 5.4).
4. Decision on the type of optimizer to be used (see Sect. 5.5).
In practice, the training data do not comprise of only pure pixels. Some of
them are mixed pixels due to either the noise present in the training data or
the mixture of classes during the selection of training data due to mislabeling
thereby introducing classification errors. A penalty parameter is introduced to
handle this error. The penalty parameter is also called the C value and depends
on the training data and the type of kernel. In general, it is found by trial and
error.
A kernel function determines the characteristic of an SVM. When the classes
in the input data are separable using a linear hyperplane, a linear kernel is used.
Similarly, a nonlinear kernel function may be used to separate data when the
classes are nonlinearly separable. However, it is very difficult to determine
which non-linear kernel function is suitable for a given application. In practice, a user will apply standard kernel functions such as a linear, polynomial,
RBF (also known as the Gaussian kernel), or sigmoid kernel. If these standard
kernel functions are unable to perform adequately, a user may have to design
a problem-specific kernel function. Moreover, the choice of the kernel further
brings in another issue that a user has to consider, which is to determine the values of parameters for a given kernel function. These parameters may be called
hyperparameters. For example, in case of an RBF kernel, the hyperparameter
that needs to be determined is a (see Sect. 5.3.3).
Thus, finding the parameters or hyperparameters that are able to accurately
predict the unknown data for a given application requires a method of model
selection or parameter search. The unknown data, in this case, refers to testing
samples or the image to be classified during the allocation stage. It is also a fact
that a model that achieves high training accuracy may not guarantee high
accuracy for the unknown data. Therefore, to establish that the model works
well on the whole dataset (i. e. the model has high generalization capacity),
a common strategy is to split the data into two parts - a training set, and
a validation set or a testing set. An SVM is trained on the training set and is
evaluated for its accuracy on the testing set. Another model selection method
is called k-fold cross-validation. In this method, the training data are divided
into k subsets of equal size. An SVM is trained on the k - 1 subsets of data
and is tested on the remaining one subset. Training and testing are performed
