Support Vector Machines
147
The concept of creating a linear separating hyperplane in a higher dimensional feature space can be incorporated in SVMs by rewriting the dual optimization problem defined in (5.23) and (5.45) as
k
1 k k
maxL (w, b,A) = LAi - - L LAiAjYiYj (Xi· Xj) .
A
. 2 .
.
l=1
l=1 J=1
(5.48)
Let a nonlinear transformation function cp map the data into a higher dimensional space. In other words, cp(x) represents the data X in the higher
dimensional space. The dual optimization problem, then, for a nonlinear case
may be expressed as
(5.49)
A kernel function is substituted for the dot product of the transformed
vectors. The explicit form of the transformation function cp is not necessarily
known. Further, the use of the kernel function is less computationally intensive.
The formulation of the kernel function from the dot product is a special case
of Mercer's theorem (Mercer 1909; Courant and Hilbert 1970; Scholkopf 1997;
Haykin 1999; Cristianini and Shawe-Taylor 2000; Scholkopf and Smola 2002).
Suppose there exists a kernel function K such that
(5.50)
The dual optimization problem for a nonlinear case can then be expressed as
k
1 k k
maxL(w,b,A) = LAi - - LLAiAjYiYjK(Xi,Xj),
A
. 2 .
l=1
l=1 j=1
subject to the constraints
k
LAiYi = 0
i=1
and
C :::: Ai :::: 0 for i = 1, ... , k .
(5.51)
(5.52)
(5.53)
In a manner similar to the other two cases, the dual optimization problem
can be solved by using Lagrange multipliers that maximizes (5.51) under the
constraints (5.52) and (5.53). The decision function can be expressed as
f(x) = sign (
L
YiAi'K (Xi, X) + b
O )
•
support vectors
(5.54)
Précédent

- 156/327

Suivant