12.1 Information, Probabilities, and Codes
195
12.1 Information, Probabilities, and Codes
Information is commonly quantified as a sequence of yes-no decisions. This leads to
the well-known binary representation with the number “1” corresponding to “yes”
and “0” corresponding to “no.” Here, each decision carries the information of one
binary digit, or bit. We therefore measure information by the number of bits needed
to uniquely describe a particular configuration among a number of equivalent configurations. Visualizing a single bit as an electrical switch, we see that the switch
describes one of two possible configurations—or micro-states. Either the switch is
on, or it is off. If both states appear equally often, with probability p = 1/2, we
can assume that simply guessing the micro-state of the switch at any time, we guess
correctly about half the time. More generally, note that the number n of possible
configurations, all assumed to have the same probability p is n = 1/ p. This even
holds if n is larger than two. As long as all of the n micro-states are equally probable,
we can ask ourselves: how many switches are needed to uniquely describe one of
the n states. The number of switches then describes the information ˆ
H needed to
identify one configuration among the n possible micro-states, which is given by
ˆ
H = log 2 (n) = log 2 (1/ p) .
(12.1)
Note that ˆ
H is the information measured in bits. This measure of information was
originally introduced by Hartley [1] in 1928, although it only became widely known
after Shannon published his seminal report [2] in 1948.
In the literature the quantity ¯
H is often referred to by “information”, “uncertainty,”
or “entropy.” Before identifying a particular one among the n possible selections, ˆ
H
refers to the “uncertainty” of our knowledge. It is also referred to as “entropy,” by
analogy to the thermodynamic entropy, which also describes our inability to identify
a particular configuration of a physical system. We will address this analogy further
in Sect. 12.2. On the other hand, once we know which one among the n different
choices is selected, we can refer to this knowledge as “information.” Usually, we can
deduce from the context which interpretation is the intended one.
The emphasis on “equally probable” in the previous paragraphs is justified by
the existence of systems, where the states are not equally probable. Consider, for
example, the ASCII [3] encoding for characters used on most computer systems. It
uses a group of eight bits to encode the normal characters and special characters,
such as comma, period, but also line feed. To identify the holders of bank accounts,
however, we only need the 26 capital letters A to Z plus a space to separate first and
family name. For this limited character set, we only need to distinguish 27 symbols
or, using the next larger power of two: five bits to distinguish 32 symbols. Moreover,
the different symbols do not occur with the same probability; the letter “E” is much
more likely than the letter “Q.” Consistent with the observation, supported by (12.1),
that events with a small probability p carry more information, we find that despite
Précédent

- 203/292

Suivant