ease the life of a system engineer of fault-tolerant systems: all described faults
should be tolerated within a limited and specified period of time.
This period actually determines the availability of the system. Fault types differ
by their impact, as well as the way they are handled.
Thus, the fault model has its own hierarchy, including single bit, element,
behavioral, and subsystem faults. One has to accept that the fault type is varying
and some action hierarchy to tolerate them is also required. All faults types should
be tolerated as there are no such systems called half- or semi-fault tolerant.
The so-called fault encapsulation approach to fault handling can help: due to
deliberate design solutions it is possible to ensure that severe faults in the system
will manifest themselves as simpler to handle faults from the system’s point of view
therefore making the fault handling practically possible to implement. This
approach will be further developed and applied here.
RT FT system applications assume long operational life; however,
fault-handling schemes are needed much more often toward the end of the device
life cycle. The appropriate techniques for tolerating faults of various types are
presented in Table 3.4. To tolerate malfunctions, time redundancy in hardware
(e.g., instruction re-execution) might be effectively used and implemented. System
software support is also needed as the hardware cannot cover all possible faults.
It is obvious that faults occurring at the bit level (stuck zero, stuck one, and
similar) should be efficiently handled ASAP (as soon as possible) and ALAP (as
local as possible), i.e., at the same or nearest level.
The term “level” in our case means the level in the hardware hierarchy on which
the fault should be handled. In other words, when a “stuck to zero” permanent fault
has happened in the register file with no corrective schemes available, the whole
register file has to be replaced, if no other possible reconfigurations were
predefined.
Fig. 3.5 New feature of
fault-tolerant system—
reliability
3.3 Models for Fault Tolerance
19
Précédent

- 35/315

Suivant