146
5: Mahesh Pal, Pakorn Watanachaturaporn
subject to the constraints
k
LAiYi = 0
(5.46)
i=1
and
C ::: Ai ::: 0 for i = 1, ... , k .
(5.47)
It can thus be seen that the objective function of the dual optimization
problem for the linearly non-separable case is the same as that of the linearly
separable case except that the Lagrange multipliers are bounded by the penalty
value C.
After obtaining the solution of (5.45), wand b can be found in the same
manner as explained in (5.26) and (5.27) earlier. The decision rule is also the
same as defined in (5.28).
5.3.3
Non-Linear Support Vector Machines
SVM seeks to find a linear separating hyperplane that can separate the classes.
There are instances where a linear hyperplane cannot separate classes without misclassification. However, those classes can be separated by a nonlinear
separating hyperplane. In fact, most of the real-life problems are non-linear
in nature (Minsky and Papert 1969). In this case, data are mapped to a higher
dimensional feature space with a nonlinear transformation function. In the
higher dimensional space, data are spread out, and a linear separating hyperplane can be constructed. For example, two classes in the input space of
Fig. 5.3 may not be separated by a linear hyperplane but a nonlinear hyperplane can make them separable. This concept is based on Cover's theorem on
the separability of patterns (Cover 1965).
Input Space
• •
.@ .
• • •
••
•
•
• •
Feature Space
• •
•
Fig.5.3. Non linear case. Mapping nonlinear data to a higher dimensional feature space
where a linear separating hyperplane can be found
Précédent

- 155/327

Suivant