288
F. Firouzi et al.
X
Y
X
Y
Big
small
X
Y
4x + 6y + (-10) = 0
8x + 12y + (-20) = 0
2x + 3y + (-5) = 0
All represent the same
line
Margin =
Margin =
Fig. 5.41 An example of hyperplanes and the impact of coefficients/weights
Fig. 5.42 An example of not
linearly separable data
distance. To prevent data points from being in the margin area, we need to add the
following constraints during the optimization:
− → w . − → x i − b ≥ 1, if y i = 1
− → w . − → x i − b < 1, if y i = −1
which can be rewritten as (for each data point)
y i
− → w . − → x i − b
≥ 1, for all 1 ≤ i ≤ n
To put it all together, we should consider minimizing
− → w
while considering the
above constraints (when the classes are linearly separable). This is an optimization
problem that can be solved by the Lagrangian multiplier method [9]. If the data are
not linearly separable (Fig. 5.42), we cannot use the above constraints during the
optimization. In this case, we need to find a hyperplane that penalizes data points
on the wrong side. To do so, we can take advantage of the hinge loss function to
penalize the points on the wrong side:
Précédent

- 294/647

Suivant