In turn, when only one register file is integrated into the chip and no other
reconfigurations are defined then the whole chip has to be replaced, etc. Pursuing
these two principles allows limiting the fault spreading and its impact to a higher
level either in the chip or the system as a whole.
Example: To tolerate bit-flip faults, hardware and system software information
redundancies might be used, as well as hardware structural support. In this sense,
parity checking in registers, supported and implemented concurrently by hardware,
is described as HW(∂I). HW(∂S) and HW(∂T) are needed as supportive
redundancies.
HW(∂S) describes the additional parity line and comparison logic, and HW(∂T)
describes the additional time needed to update the parity line and executing the
comparison. However, the main type of redundancy used in this approach is
information.
Up to the best knowledge of the authors, there have been no representative
statistics which characterize the exact distribution of faults for computer systems.
The distribution of faults depends on the operational environment, for example,
temperature, vibration, and radiation exposure.
Even so, it is a well-known fact that the ratio of malfunction to permanent faults
can be up to 10
3
–10
6
. The upper bound belongs to aerospace and aviation, principally due to malfunctions, i.e., errors induced by alpha particles.
In this sense, Figs. 3.4 and 3.5 are transformed into Fig. 3.6 which presents
various faults in the system and various possible solutions. M fault illustrates the fact
that the fault types are not separated. For example, Byzantine faults of the system
might be “stuck at zero” faults of the hardware that were spread throughout the
system.
Table 3.4 Classification of HW faults in FT computer systems
Type of
Fault
Description
Impact
Byzantine
An arbitrary behavior of a part of a
device, hardware, or a program
The entire system is affected
Subsystem
faults
An arbitrary behavior of a subsystem
of the processor temporally or
permanently
The entire system is affected
Open fault
Resistance on either a line or a block
occurs due to a bad connection
The value associated with the line or
the block is modified
Bridging
fault
Crossing lines, the number of lines
crossed varies
The value associated with the line or
the block is modified to a different
value
Stuck at 0,
Stuck at 1
The result value is fixed to 0 or 1
The result value is stuck to 0 or 1
Bit-flip
fault
The result value is fixed to 0 or 1
The bit is modified
20
3 Fault Tolerance: Theory and Concepts
reconfigurations are defined then the whole chip has to be replaced, etc. Pursuing
these two principles allows limiting the fault spreading and its impact to a higher
level either in the chip or the system as a whole.
Example: To tolerate bit-flip faults, hardware and system software information
redundancies might be used, as well as hardware structural support. In this sense,
parity checking in registers, supported and implemented concurrently by hardware,
is described as HW(∂I). HW(∂S) and HW(∂T) are needed as supportive
redundancies.
HW(∂S) describes the additional parity line and comparison logic, and HW(∂T)
describes the additional time needed to update the parity line and executing the
comparison. However, the main type of redundancy used in this approach is
information.
Up to the best knowledge of the authors, there have been no representative
statistics which characterize the exact distribution of faults for computer systems.
The distribution of faults depends on the operational environment, for example,
temperature, vibration, and radiation exposure.
Even so, it is a well-known fact that the ratio of malfunction to permanent faults
can be up to 10
3
–10
6
. The upper bound belongs to aerospace and aviation, principally due to malfunctions, i.e., errors induced by alpha particles.
In this sense, Figs. 3.4 and 3.5 are transformed into Fig. 3.6 which presents
various faults in the system and various possible solutions. M fault illustrates the fact
that the fault types are not separated. For example, Byzantine faults of the system
might be “stuck at zero” faults of the hardware that were spread throughout the
system.
Table 3.4 Classification of HW faults in FT computer systems
Type of
Fault
Description
Impact
Byzantine
An arbitrary behavior of a part of a
device, hardware, or a program
The entire system is affected
Subsystem
faults
An arbitrary behavior of a subsystem
of the processor temporally or
permanently
The entire system is affected
Open fault
Resistance on either a line or a block
occurs due to a bad connection
The value associated with the line or
the block is modified
Bridging
fault
Crossing lines, the number of lines
crossed varies
The value associated with the line or
the block is modified to a different
value
Stuck at 0,
Stuck at 1
The result value is fixed to 0 or 1
The result value is stuck to 0 or 1
Bit-flip
fault
The result value is fixed to 0 or 1
The bit is modified
20
3 Fault Tolerance: Theory and Concepts
