Support Vector Machines
139
controlled by the slack variables and the regularization parameter. The VCdimension is minimized in the same fashion as in the linearly separable case.
Often, a linear separating hyperplane is not able to classify input data (either
noiseless or noisy), but a non-linear separating hyperplane can. This has been
referred to as the nonlinear case (Boser et al. 1992). In this case, the input
data are transformed into a higher dimensional space that spreads the data
out such that a linearly separable hyperplane can be obtained. They also
suggested the use of kernel functions as transformation functions to reduce
the computational cost. The design of SVMs for the three cases is described
next.
5.3.1
Linearly Separable Case
Linearly separable case is the simplest of all to design a support vector machine.
Consider a binary classification problem under the assumption that data can
be separated into two classes using a linear separating hyperplane (Fig. 5.1).
Consider k training samples obtained from the two classes, represented by
(Xl> yI) , ... , (Xb Yk), where Xj E ]RN is an N-dimensional observed data vector
with each sample belonging to either of the two classes labeled by y E {-I, + I}.
These training samples are said to be linearly separable if there exists an Ndimensional vector w that determines the orientation of a discriminating plane
and a scalar b that determines the offset of this plane from the origin such that
w . Xi + b ::: + 1 for all y = + 1 ,
W . Xi + b ::: -1 for all y = -1 .
(s.8)
(s.9)
Points lying on the optimal hyperplane satisfy the equation WXj + b = o. The
Feature 1
Fig.5.1. A linear separating hyperplane for the linearly separable data sets. Dashed lines
pass through the support vectors
Précédent

- 148/327

Suivant