146
B.G.H. Gorte
P (Ci ) : prior probability
P(Ci) is called the prior probability of class Ci , the probability for any pixel
that it belongs to Ci , irrespective of its feature vector. It can be estimated on the
basis of prior knowledge about the terrain, as the (relative) area that is expected
to be covered by Ci .
P(x) : (class-independent) feature probability density
For a certain x, in order to find the class Ci with the maximum posterior probability P(Cilx) the various P(Cilx) must be only compared. This comparison
can be done with the same success without knowing P(x) since it is the same
for every Ci . Therefore, it is common to substitute P(x) by a normalization
factor
N
P(x) = LP(xICj) P(Cj )
(7.5)
j=1
This leaves, however, no room for an unknown class, since there is no way to
find P(xlunknown) (by definition of unknown).
Gaussian Maximum Likelihood. To estimate class probability desities, it is commom to assume a multivariate normal (Gaussian) distribution of P(xICi) for each
class Ci . . From the training data, for each class Ci the sample mean vector and the
sample variance-covariance matrix are obtained, and they are used as estimates for
the class mean vector mi and the class variance- covariance matrix Vi, This gives
the class probability density function for class Ci :
with:
M : the number of features
Vi : the M x M variance-covariance matrix of class Ci
IViI : the determinant of Vi
Vi-I: the inverse of Vi
(7.6)
y : x - mi (mi is the class mean vector), as a column vector with M components
yT : the transposed ofy (a row vector).
A classifier that does not take prior probabilities into account only needs probability densities P(xICi)for different Ci . By removing constants, taking a logarithms
and multiplying the result by -2, equation 7.6 can be written as:
(7.7)
Di is a function of x and can be regarded as a measure of the dissimilarity between a feature vector x and a class mean mi (because of y T and y) in the feature
space. It is weighted by Vi-I and compensated by In IViI according to the withinclass variability: From a given x, a widely spread class seems more similar than
a concentrated one. The second term of equation 7.7 is commonly referred to as
Mahalanobis distance
B.G.H. Gorte
P (Ci ) : prior probability
P(Ci) is called the prior probability of class Ci , the probability for any pixel
that it belongs to Ci , irrespective of its feature vector. It can be estimated on the
basis of prior knowledge about the terrain, as the (relative) area that is expected
to be covered by Ci .
P(x) : (class-independent) feature probability density
For a certain x, in order to find the class Ci with the maximum posterior probability P(Cilx) the various P(Cilx) must be only compared. This comparison
can be done with the same success without knowing P(x) since it is the same
for every Ci . Therefore, it is common to substitute P(x) by a normalization
factor
N
P(x) = LP(xICj) P(Cj )
(7.5)
j=1
This leaves, however, no room for an unknown class, since there is no way to
find P(xlunknown) (by definition of unknown).
Gaussian Maximum Likelihood. To estimate class probability desities, it is commom to assume a multivariate normal (Gaussian) distribution of P(xICi) for each
class Ci . . From the training data, for each class Ci the sample mean vector and the
sample variance-covariance matrix are obtained, and they are used as estimates for
the class mean vector mi and the class variance- covariance matrix Vi, This gives
the class probability density function for class Ci :
with:
M : the number of features
Vi : the M x M variance-covariance matrix of class Ci
IViI : the determinant of Vi
Vi-I: the inverse of Vi
(7.6)
y : x - mi (mi is the class mean vector), as a column vector with M components
yT : the transposed ofy (a row vector).
A classifier that does not take prior probabilities into account only needs probability densities P(xICi)for different Ci . By removing constants, taking a logarithms
and multiplying the result by -2, equation 7.6 can be written as:
(7.7)
Di is a function of x and can be regarded as a measure of the dissimilarity between a feature vector x and a class mean mi (because of y T and y) in the feature
space. It is weighted by Vi-I and compensated by In IViI according to the withinclass variability: From a given x, a widely spread class seems more similar than
a concentrated one. The second term of equation 7.7 is commonly referred to as
Mahalanobis distance
