136
Biomedical Signal and Image Processing
The simple Bayesian classification method, although straightforward and
intuitive, may not be sophisticated enough to deal with real-world problems. In
practice, simple Bayesian methods are often combined with some heuristic and
domain-based knowledge to improve the classification results. One extension of the
Bayesian classification method is achieved through the use of the concept of “loss
function.”
7.5.1 LOSS FUNCTION
In many practical applications including medical diagnostics, it is often the case
that some misclassifications are more costly than others. For instance, the cost of
misclassifying a cancer sample as normal (healthy) is far more than misclassifying
a normal sample as cancer. As a result, one would like to somehow incorporate the
importance and impact of each decision directly in the classification algorithm.
Loss functions are used to give different weights to different classification mistakes. Assuming the two-category classification, such as in the simple example
discussed earlier, there will be two types of acts: α 1 and α 2 . Now, let us define
λ(α i |ω j ) as the loss when action α i is taken, while the true state of the nature is ω j .
Using the loss functions λ(α i |ω j ), i = 1, 2, j = 1, 2 (i ≠ j), a new risk function for each
action can be defined to improve the decision criteria and incorporate the relative
importance of each misclassification in the decision process. This risk function
can be defined as follows:
a
l a w P w
l a w P
R( i |X) = ( i | 1 ) ( 1 |X) + ( i | 2 ) ( w 2 |X)
(7.7)
The best decision is then made based on these risk functions. For instance, if
R(α 1 |X) < R(α 2 |X), α 1 will be the action.
Example 7.4
Figure 7.4 shows two categories forming a 2-D dataset. The data for each category
are Gaussian distributed. We intend to use Bayesian decision rule to find a decision
boundary between the two categories.
To do this, we need to estimate probability function for each category. Since
we know that probability functions of these two categories are Gaussian, all we
need to do is to compute the sample mean and sample variance of each category
to obtain the probability density functions (PDFs). Calculating the sample means
and sample variances of these two categories, we have
4⎤
1⎤
1 2
0 ⎤
/
0 ⎤
⎡
⎡−
⎡ /
⎡1 2
m 1 =
m 2 =
Σ 1 =
Σ 2 =
⎢ ⎥
⎢ ⎥
⎢
⎥
⎢
⎥
4
4
0
1 2
0
1 2
/
/
⎣ ⎦
⎣ ⎦
⎣
⎦
⎣
⎦
Using Equation 7.4, the optimal decision boundary can be obtained as follows:
1 ) ( 1 P X w 2 w 2 )
(7.8)
P X
( |w P w ) = ( | ) (
P
Précédent

- 163/412

Suivant