138
Biomedical Signal and Image Processing
x 1
w 1
1
b
y
x 2
w 2
x n
w n
FIGURE 7.5 Structure of a perceptron.
7.6 MAXIMUM LIKELIHOOD METHOD
One of the main problems in pattern recognition is parameter estimation of the
probability functions. Maximum likelihood estimation (MLE) is one of the most
popular methods to address this problem. Methods based on MLEs often provide
better convergence properties as the number of training samples increases.
As mentioned previously, in order to form the Bayesian decision equations,
the class conditional probabilities are needed. Since in practice, the knowledge
of these probabilities is not available, efficient techniques are employed to estimate these probabilities. Assume that the type or family of the PDF p(X|ω i ), for
example, Gaussian, is known, one can determine the parameters of this probability
function.
Suppose that there are c classes in a given pattern classification problem. Also,
assume that the n sample x 1 ,…, x n created by one of the classes (according to the PDF
of that class) are collected in a set D. Note that we do not know which class this set
of observed data belongs to and we are to use the MLE to predict which one of these
classes c is more likely to have generated the data. We assume that the samples in
D are independent and identically distributed (i.i.d.). As mentioned before, we are
assuming that the PDFs, i.e., p(X|θ i )’s for i = 1, 2,…, c, have a known parametric
form, and, therefore, we only need to determine the parameter vector θ that uniquely
identifies the PDF of one of the classes. Then, knowing that the samples are independent, we can describe the probability of obtaining the sample set D assuming that the
parameters are θ as follows:
n
p D| )
( q ∏
=
=
k 1
p X |q )
(
(7.12)
k
We call p(D|θ) the likelihood of θ because this probability identifies the likelihood
of obtaining the samples in D given that the estimated parameter set θ is a suitable
one. The MLE is simply an estimation method that finds the set of parameters that
maximizes the likelihood function p(D|θ). In other words, the maximum likelihood
function results in a probability function that is more likely to produce the observed
data. Note that since we do not know the true set of parameters, all we can obtain
is an estimation of these values through methods such as maximum likelihood.
As such, we denote by θ ˆ the MLE of θ, which maximizes p(D|θ).
Précédent

- 165/412

Suivant