Support Vector Machines
149
5.4.1
One Against the Rest Classification
This method is also called winner-take-all classification. Suppose the dataset
is to be classified into M classes. Therefore, M binary SVM classifiers may
be created where each classifier is trained to distinguish one class from the
remaining M - 1 classes. For example, class one binary classifier is designed
to discriminate between class one data vectors and the data vectors of the
remaining classes. Other SVM classifiers are constructed in the same manner.
During the testing or application phase, data vectors are classified by finding
the margin from the linear separating hyperplane 0. e. (5.28) or (5.54) without
the sign function}:
m
i(x} = LYiA1K (x, Xi) + lJ for j = 1, . .. ,M ,
(5.55)
i=1
where m is the number of support vectors and M is the number of classes.
A data vector is assigned the class label of the SVM classifier that produces the
maximum output.
However, if the outputs corresponding to two or more classes are very close
to each other, those points are labeled as unclassified, and a subjective decision
may have to be taken by the analyst. Otherwise, a reject decision (Scholkopf
and Smola 2002) may also be applied using a threshold to decide the class
label. For example, considering the two largest outputs resulting from (5.55),
the difference between the two outputs must be larger than the reject decision
threshold to assign a class label.
This multiclass method has an advantage in the sense that the number of
binary classifiers to construct equals the number of classes. However, there are
some drawbacks. First, during the training phase, the memory requirement
is very high and is proportional to the square of the total number of training
samples. This may cause problems for large training data sets and may lead
to computer memory problems. Second, suppose there are M classes and each
has an equal number of training samples. During the training phase, the ratio
of training samples of one class to the rest of the classes will be 1 : (M - I). This
ratio, therefore, shows that training sample sizes will be unbalanced. Because
of these limitations, the pairwise multiclass method has been proposed.
5.4.2
Pairwise Classification
In this method, SVM classifiers for all possible pairs of classes are created
(Knerr et al. 1990; Friedman, 1996; Hastie and Tibshirani 1998; KreBeI1999).
Therefore, for M classes, there will be ! M(M - 1) binary classifiers. The output
from each classifier in the form of a class label is obtained. The class label that
occurs the most is assigned to that point in the data vector. In case of a tie,
a tie-breaking strategy may be adopted. A common tie-breaking strategy is to
randomly select one of the class labels that are tied.
149
5.4.1
One Against the Rest Classification
This method is also called winner-take-all classification. Suppose the dataset
is to be classified into M classes. Therefore, M binary SVM classifiers may
be created where each classifier is trained to distinguish one class from the
remaining M - 1 classes. For example, class one binary classifier is designed
to discriminate between class one data vectors and the data vectors of the
remaining classes. Other SVM classifiers are constructed in the same manner.
During the testing or application phase, data vectors are classified by finding
the margin from the linear separating hyperplane 0. e. (5.28) or (5.54) without
the sign function}:
m
i(x} = LYiA1K (x, Xi) + lJ for j = 1, . .. ,M ,
(5.55)
i=1
where m is the number of support vectors and M is the number of classes.
A data vector is assigned the class label of the SVM classifier that produces the
maximum output.
However, if the outputs corresponding to two or more classes are very close
to each other, those points are labeled as unclassified, and a subjective decision
may have to be taken by the analyst. Otherwise, a reject decision (Scholkopf
and Smola 2002) may also be applied using a threshold to decide the class
label. For example, considering the two largest outputs resulting from (5.55),
the difference between the two outputs must be larger than the reject decision
threshold to assign a class label.
This multiclass method has an advantage in the sense that the number of
binary classifiers to construct equals the number of classes. However, there are
some drawbacks. First, during the training phase, the memory requirement
is very high and is proportional to the square of the total number of training
samples. This may cause problems for large training data sets and may lead
to computer memory problems. Second, suppose there are M classes and each
has an equal number of training samples. During the training phase, the ratio
of training samples of one class to the rest of the classes will be 1 : (M - I). This
ratio, therefore, shows that training sample sizes will be unbalanced. Because
of these limitations, the pairwise multiclass method has been proposed.
5.4.2
Pairwise Classification
In this method, SVM classifiers for all possible pairs of classes are created
(Knerr et al. 1990; Friedman, 1996; Hastie and Tibshirani 1998; KreBeI1999).
Therefore, for M classes, there will be ! M(M - 1) binary classifiers. The output
from each classifier in the form of a class label is obtained. The class label that
occurs the most is assigned to that point in the data vector. In case of a tie,
a tie-breaking strategy may be adopted. A common tie-breaking strategy is to
randomly select one of the class labels that are tied.
