240
10: Pakorn Watanachaturaporn, Manoj K. Arora
on the input data k times with different k - 1 subsets for training and the
remaining one subset for testing. In other words, the testing subset is used to
assess the accuracy of the model trained by the k -1 subsets. After repeating the
above process k times, the overall accuracy measured by the cross-validation
method is the percentage of the number of correctly classified data for all k
testing subsets. The model that produces the highest overall accuracy is used
to classify data during the allocation stage.
A straightforward way to find a judicious combination of the penalty value
and the hyperparameters of kernel functions is to use a grid-search method.
In the grid-search method, various combinations of the penalty value and
the hyperparameters are formed and their effect on the accuracy is assessed
by using either a validation set or k-fold cross-validation. The model that
gives the best accuracy is selected. Applying the grid-search method is very
time consuming because it acts on every combination of the penalty value
and hyperparameters. To reduce the model selection time, one may perform
a coarse grid-search first. For example, the penalty value may take values 10 2
apart such as 10- 5 , 10- 3 , .•• , 105. Then, after identifying the best classification
accuracy region obtained from the coarse grid-search, a finer grid-search may
be conducted over that region only. For instance, the best accuracy from the
coarse grid search might be between the penalty values of 10 1 to 10 3 . The finer
grid search might be performed over penalty values with finer resolution (e. g.
10 1 ,101.1, ... ,10 3 ). The final model is trained using the whole training set with
the hyperparameters identified by the grid-search method.
Another issue related to SVM classifiers is the selection of appropriate multiclass method (see Sect. 5.4). SVMs were initially developed to perform binary
classification; though, applications of binary classification are very limited.
More practical applications involve multiclass classification. For example, in
remote sensing, land cover classification is generally a multiclass problem.
A number of methods have been proposed to employ SVMs to produce multiclass classification. Most of the methods generate multiclass classifications
from a number of binary SVM classifiers. All the methods have their own merits
and demerits. For example, the one against the rest approach gives an advantage in terms of simplicity and requires fewer binary classifiers. However, many
have argued that in this method the number of training samples of each class for
each binary classifier become significantly imbalanced. To solve this problem,
methods such as a re-sampling method have been proposed. In the re-sampling
method, pseudo training data are created by copying existing training data and
adding/subtracting noise. In other words, a new training area is identified for
a class and a number of training data are selected randomly within that area.
The latter approach may help in improving the accuracy in some cases. However, it results in increased training time since the number of training data
is increased due to the use of the pseudo training data. Both the approaches
are simpler than the pairwise and directed acyclic graph (DAG) approaches.
However, experiments show that, training time of the one against the rest both
with and without balancing training data is longer and percent accuracy is
lower. On the contrary, the pairwise and DAG approaches require more binary
classifiers, but they take less training time and give more accurate results.
Précédent

- 247/327

Suivant