115
Other Signal and Image Processing Methods
for an outcome with p i = 0 is still infinity. In order to address this issue, we can
further modify our measure to “p i log(1/p i ).” This measure gives zero information for
outcomes with zero probability.
While the measure defined as “p i log(1/p i )” correctly identifies the information
of “each outcome,” it fails to give an average amount of information for the random
variable X as a whole. A logical approach to define such a measure is to add up the
information measures of all N outcomes. This leads to the following measure, commonly referred to as entropy:
N 1
i
−
∑
⎛
⎞

1

H X
( ) =

p i log
(6.31)

⎜
⎝

⎟
⎠

p i
0
=
The base in the preceding log function can be any number, but the most popular
bases are “2,” “e,” and “10.” Let us consider a random variable that is reduced to
a deterministic signal, i.e., when for some i = k: p i = 0 and for all other i’s: p i = 0.
For such a case, from the earlier definition, we have H(X) = 0. This simply means
that if one knows the outcome of a process beforehand (i.e., the k’s outcome always
happens), there is no surprise and therefore no information in the process. This is
expected from a suitable measure of information, as discussed before.
The second observation deals with the other extreme case, i.e., when every outcome has the same probability. In such a case, we have p i = 1/N, and, therefore,
entropy is H(X) = log(N). One can easily prove that this is maximum possible
entropy. In such random processes, we have no bias or guess before the outcome of
the variable is identified. For instance, tossing a fair dice gives in general maximum
information because no one can guess what number is going to show up on the dice
beforehand, and, therefore, any outcome is very informative.
There are many other secondary definitions based on the basic idea of entropy.
One of these measures is conditional entropy. Consider two random variables X and Y.
Assume that we know the outcome of X statistically, i.e., based on the probabilities
of the outcomes in X, we can expect what outcomes are happening more and what
outcomes are less likely to happen. Knowing such information about X, now we want
to discover how much information is gained when Y is revealed to us. For instance,
if we already know that a patient has had a myocardial infarction (heart attack), how
surprised we will be to know that there is a significant change in patient’s blood
pressure. Obviously, since we know that a heart attack has happened we are less
surprised to hear that there have been some fluctuations in the blood pressure. The
conditional entropy of Y given X identifies in average how much surprise is left in
knowing Y, given that X has already happened. This measure is defined as follows:
−
−
∑
1 ∑
1
⎛
⎞

N
M
1

p j i
⎜
⎜
⎝

⎟
⎟
⎠

H Y X )
( |
=

p i j
, log
2
(6.32)

i=0 i=0
In the preceding equation, p i,j is the joint probability of the variables X and Y, and p j|i
is the conditional probability of the variable Y given X.
Précédent

- 142/412

Suivant