5 Machine Learning for IoT
289
Big Margin | Classification Error
Small Margin | Correct Classification
Fig. 5.43 Tradeoff between classification error and margin
X
Y
Y
X
Z
X
a)
b)
c)
Fig. 5.44 Kernel trick in SVM. (a) Non-linearly separable data. (b) Data on higher dimension and
a linear decision boundary. (c) Decision boundary in original dimensions
max
0, 1 − y i
− → w . − → x i − b
This cost function is 0 if the actual and predicted values are on the same side, but
for the wrong side points, the function’s value (penalty) is increased proportionally
to the distance of the wrong data points from the hyperplane. Now let us study the
example of Fig. 5.43. As shown in this figure, we extract a hyperplane that has a
big margin, but it cannot classify all the data points. On the other hand, we can have
a hyperplane that has a very small margin (which is not good for generalization),
but it can classify all the data points. In SVM, we can make a tradeoff between
classification error and margin. To do so, we define a regularization parameter called
C and the SVM cost function is defined as
Cost = max
0, 1 − y i
− → w . − → x i − b
+ C ∗
− → w
In a case that data cannot be classified by a linear hyperplane (similar to
Fig. 5.44a), we can apply a technique called kernel trick. The idea is to use a
nonlinear function (kernel) to map the data points to a new high dimensional space,
in which we can find a linear hyperplane using the abovementioned linear SVM
technique [9, 10]. Let us apply the kernel trick to our example (Fig. 5.44a) to clarify
Précédent

- 295/647

Suivant