136
5: Mahesh Pal, Pakorn Watanachaturaporn
may not be able to evaluate the integral in (5.1). The only known information
is contained in the training samples. Therefore, a stochastic approximation of
integral in (5.1) is desired that can be computed empirically by a finite sum
given by
(5.2)
and, thus, is known as the empirical risk. The value Remp (a) is a fixed number for
a given a and a particular training data set. Next we discuss the minimization
of risk functions.
5.2.1
Empirical Risk Minimization
The empirical risk is different from the expected risk in two ways (Haykin
1999):
1. it does not depend on the unknown cumulative distribution function
2. it can be minimized with respect to the parameter a
Based on the law of large numbers (Gray and Davisson 1986), the empirical
mean of a random variable converges to its expected value if the size of the
training samples is infinitely large. This remark justifies the use of the empirical
risk Remp(a) instead of the risk function R(a). However, convergence of the
empirical mean of the random variable to its expected value does not imply
that the value a that minimizes the empirical risk will also minimize the
risk function R(a). If convergence of the minimum of the empirical risk to
the minimum of the expected risk does not occur, this principle of empirical
risk minimization is said to be inconsistent. In this case, even though the
empirical risk is minimized, the expected risk may be high. In other words,
a small error rate of a learning machine on the training samples does not
necessarily guarantee high generalization ability (i. e. the ability to work well
on unseen data). This situation is commonly referred to as overfitting. Vapnik
and Chervonenkis (1971, 1991) have shown that consistency occurs if and
only if convergence in probability of the empirical risk to the expected risk
is substituted by uniform convergence in probability. Note that convergence in
probability of R(a) means that for any E > 0 and for any 'l > 0, there exists
a number ko = ko (E, 'l) such that for any k > ko, the inequality R (ak) - R (ao) <
E holds true with a probability of at least 1-'l (Vapnik 1999). Eis a small number
close to zero, and 'l is referred to as the level of significance - similar to the a
value in statistics. Uniform convergence in probability is defined as
lim Prob (sup IR(a) - Remp(a) I > E) -+ 0, VE,
k---+oo
aEA
(5.3)
Précédent

- 145/327

Suivant