38
Digital Electronics
The Code Table
The code table defined by both Unicode and ISO-10646 provides a unique number for every character,
irrespective of the platform, program and language used. The table contains characters required to
represent practically all known languages and scripts. The list includes not only the Greek, Latin,
Cyrillic, Arabic, Arabian and Georgian scripts but also Japanese, Chinese and Korean scripts. In
addition, the list also includes scripts such as Devanagari, Bengali, Gurmukhi, Gujarati, Oriya, Telugu,
Tamil, Kannada, Thai, Tibetan, Ethiopic, Sinhala, Canadian Syllabics, Mongolian, Myanmar and
others. Scripts not yet covered will eventually be added. The code table also covers a large number of
graphical, typographical, mathematical and scientific symbols.
In the 32-bit version, which is the most recent version, the code table is divided into 2
16 subsets, with
each subset having 2
16 characters. In the 32-bit representation, elements of different subsets therefore
differ only in the 16 least significant bits. Each of these subsets is known as a plane. Plane 0, called the
basic multilingual plane (BMP), defined by 00000000 to 0000FFFF, contains all most commonly used
characters including all those found in major older encoding standards. Another subset of 2
16 characters
could be defined by 00010000 to 0001FFFF. Further, there are different slots allocated within the
BMP to different scripts. For example, the basic Latin character set is encoded in the range 0000 to
007F. Characters added to the code table outside the 16-bit BMP are mostly for specialist applications
such as historic scripts and scientific notation. There are indications that there may never be characters
assigned outside the code space defined by 00000000 to 0010FFFF, which provides space for a little
over 1 million additional characters.
Different characters in Unicode are represented by a hexadecimal number preceded by ‘U+’. For
example, ‘A’ and ‘e’ in basic Latin are respectively represented by U+0041 and U+0065. The first
256 code numbers in Unicode are compatible with the seven-bit ASCII-code and its eight-bit variant
ISO-8859-1. Unicode characters U+0000 to U+007F (128 characters) are identical to those in the
ASCII code, and the Unicode characters in the range U+0000 to U+00FF (256 characters) are identical
to ISO-8859-1.
Use of Combining Characters
Unicode assigns code numbers to combining characters, which are not full characters by themselves
but accents or other diacritical marks added to the previous character. This makes it possible to place
any accent on any character. Although Unicode allows the use of combining characters, it also assigns
separate codes to commonly used accented characters known as precomposed characters. This is done
to ensure backwards compatibility with older encodings. As an example, the character ‘ä’ can be
represented as the precomposed character U+00E4. It can also be represented in Unicode as U+0061
(Latin lower-case letter ‘a’) followed by U+00A8 (combining character ‘..’).
Unicode and ISO-10646 Comparison
Although Unicode and ISO-10646 have identical code tables, Unicode offers many more features not
available with ISO-10646. While the ISO-10646 standard is not much more than a comprehensive
character set, the Unicode standard includes a number of other related features such as character
properties and algorithms for text normalization and handling of bidirectional text to ensure correct
display of mixed texts containing both right-to-left and left-to-right scripts.
2.5 Seven-segment Display Code
Seven-segment displays [Fig. 2.1(a)] are very common and are found almost everywhere, from pocket
calculators, digital clocks and electronic test equipment to petrol pumps. A single seven-segment
display or a stack of such displays invariably meets our display requirement. There are both LED and
Digital Electronics
The Code Table
The code table defined by both Unicode and ISO-10646 provides a unique number for every character,
irrespective of the platform, program and language used. The table contains characters required to
represent practically all known languages and scripts. The list includes not only the Greek, Latin,
Cyrillic, Arabic, Arabian and Georgian scripts but also Japanese, Chinese and Korean scripts. In
addition, the list also includes scripts such as Devanagari, Bengali, Gurmukhi, Gujarati, Oriya, Telugu,
Tamil, Kannada, Thai, Tibetan, Ethiopic, Sinhala, Canadian Syllabics, Mongolian, Myanmar and
others. Scripts not yet covered will eventually be added. The code table also covers a large number of
graphical, typographical, mathematical and scientific symbols.
In the 32-bit version, which is the most recent version, the code table is divided into 2
16 subsets, with
each subset having 2
16 characters. In the 32-bit representation, elements of different subsets therefore
differ only in the 16 least significant bits. Each of these subsets is known as a plane. Plane 0, called the
basic multilingual plane (BMP), defined by 00000000 to 0000FFFF, contains all most commonly used
characters including all those found in major older encoding standards. Another subset of 2
16 characters
could be defined by 00010000 to 0001FFFF. Further, there are different slots allocated within the
BMP to different scripts. For example, the basic Latin character set is encoded in the range 0000 to
007F. Characters added to the code table outside the 16-bit BMP are mostly for specialist applications
such as historic scripts and scientific notation. There are indications that there may never be characters
assigned outside the code space defined by 00000000 to 0010FFFF, which provides space for a little
over 1 million additional characters.
Different characters in Unicode are represented by a hexadecimal number preceded by ‘U+’. For
example, ‘A’ and ‘e’ in basic Latin are respectively represented by U+0041 and U+0065. The first
256 code numbers in Unicode are compatible with the seven-bit ASCII-code and its eight-bit variant
ISO-8859-1. Unicode characters U+0000 to U+007F (128 characters) are identical to those in the
ASCII code, and the Unicode characters in the range U+0000 to U+00FF (256 characters) are identical
to ISO-8859-1.
Use of Combining Characters
Unicode assigns code numbers to combining characters, which are not full characters by themselves
but accents or other diacritical marks added to the previous character. This makes it possible to place
any accent on any character. Although Unicode allows the use of combining characters, it also assigns
separate codes to commonly used accented characters known as precomposed characters. This is done
to ensure backwards compatibility with older encodings. As an example, the character ‘ä’ can be
represented as the precomposed character U+00E4. It can also be represented in Unicode as U+0061
(Latin lower-case letter ‘a’) followed by U+00A8 (combining character ‘..’).
Unicode and ISO-10646 Comparison
Although Unicode and ISO-10646 have identical code tables, Unicode offers many more features not
available with ISO-10646. While the ISO-10646 standard is not much more than a comprehensive
character set, the Unicode standard includes a number of other related features such as character
properties and algorithms for text normalization and handling of bidirectional text to ensure correct
display of mixed texts containing both right-to-left and left-to-right scripts.
2.5 Seven-segment Display Code
Seven-segment displays [Fig. 2.1(a)] are very common and are found almost everywhere, from pocket
calculators, digital clocks and electronic test equipment to petrol pumps. A single seven-segment
display or a stack of such displays invariably meets our display requirement. There are both LED and
