144
5: Mahesh Pal, Pakorn Watanachaturaporn
Feature 1
Fig. 5.2. Illustration of the linearly non-separable case
The constraint in (5.10) is thus relaxed to
Yi (W . Xi + b) - 1 + {ii ~ 0 .
(5.31)
In the objective function, a new term called the penalty value 0 < C < 00
is added. The penalty value is a form of regularization parameter and defines
the trade-off between the number of noisy training samples and the classifier
complexity. It is usually selected by trial and error.
It can be shown that the optimization problem for the linearly non-separable
case becomes
subject to the constraints
Yi (w . Xi + b) - 1 + {ii ~ 0
and
{ii ~ 0 for i = 1, ... , k .
(5.32)
(5.33)
(5.34)
The penalty value may have a significant effect on the performance of the
resulting support vector machines (see experimental results in Chap. 10).
From (5.32), it can also be seen that when C ~ 0, the minimization problem is
not affected by the misclassifications even though {ii > o. The linear separating
hyperplane will be located at the midpoint of the two classes with the largest
possible separation. When C > 0, the minimization problem is affected by {ii.
When C ~ 00, the values of {ii approach zero and the minimization problem
reduces to the linearly separable case.
5: Mahesh Pal, Pakorn Watanachaturaporn
Feature 1
Fig. 5.2. Illustration of the linearly non-separable case
The constraint in (5.10) is thus relaxed to
Yi (W . Xi + b) - 1 + {ii ~ 0 .
(5.31)
In the objective function, a new term called the penalty value 0 < C < 00
is added. The penalty value is a form of regularization parameter and defines
the trade-off between the number of noisy training samples and the classifier
complexity. It is usually selected by trial and error.
It can be shown that the optimization problem for the linearly non-separable
case becomes
subject to the constraints
Yi (w . Xi + b) - 1 + {ii ~ 0
and
{ii ~ 0 for i = 1, ... , k .
(5.32)
(5.33)
(5.34)
The penalty value may have a significant effect on the performance of the
resulting support vector machines (see experimental results in Chap. 10).
From (5.32), it can also be seen that when C ~ 0, the minimization problem is
not affected by the misclassifications even though {ii > o. The linear separating
hyperplane will be located at the midpoint of the two classes with the largest
possible separation. When C > 0, the minimization problem is affected by {ii.
When C ~ 00, the values of {ii approach zero and the minimization problem
reduces to the linearly separable case.
