At this level, only hardware-based redundancy HW(I, S, T) can be used as
software support would need too much time. Small time redundancy T is used as
supportive redundancy for less than instruction time to achieve fault tolerance, or as
redundancy HW(2T) in case of instruction repetition by the processor.
In another example, triplicate memory with voter, the voter masks the faults of
one memory element while accessing the memory and thus the fault is tolerated
within one instruction. The techniques such as microinstruction or instruction
repetition can be used here as well to provide instruction-level fault correction
within the processor.
We call this method of recovery a “good” Fault-Tolerant Computing System
(FCTS) as the fault checking, detection, and recovery are done completely transparent to the software [9]. In case of permanent faults, the system can reconfigure
itself and continue processing without application intervention.
Note that not all faults are detectable and recoverable within the instruction level,
thus the other levels of recovery are still required and should be considered as well
in design of fault-tolerant systems.
For the procedure level of recovery, a hardware fault and its influence is tolerated and eliminated within the scope of a procedure. For example, techniques
such as recovery blocks, wrongly known as recovery points [10], can be considered
as a procedure-level scheme.
In contrast to the recovery on the instruction level, the software state space must
be in this case conserved and regenerated to recover from a fault. For this scheme,
the state of the system that is required to save and recover is much larger than in the
instruction-level approach and a significant software overhead is expected to
implement fault tolerance at this level. We call a system that uses only this
mechanism therefore a “medium” FCTS.
Fig. 4.3 System recovery times according to the used scheme
34
4 Generalized Algorithm of Fault Tolerance (GAFT)
software support would need too much time. Small time redundancy T is used as
supportive redundancy for less than instruction time to achieve fault tolerance, or as
redundancy HW(2T) in case of instruction repetition by the processor.
In another example, triplicate memory with voter, the voter masks the faults of
one memory element while accessing the memory and thus the fault is tolerated
within one instruction. The techniques such as microinstruction or instruction
repetition can be used here as well to provide instruction-level fault correction
within the processor.
We call this method of recovery a “good” Fault-Tolerant Computing System
(FCTS) as the fault checking, detection, and recovery are done completely transparent to the software [9]. In case of permanent faults, the system can reconfigure
itself and continue processing without application intervention.
Note that not all faults are detectable and recoverable within the instruction level,
thus the other levels of recovery are still required and should be considered as well
in design of fault-tolerant systems.
For the procedure level of recovery, a hardware fault and its influence is tolerated and eliminated within the scope of a procedure. For example, techniques
such as recovery blocks, wrongly known as recovery points [10], can be considered
as a procedure-level scheme.
In contrast to the recovery on the instruction level, the software state space must
be in this case conserved and regenerated to recover from a fault. For this scheme,
the state of the system that is required to save and recover is much larger than in the
instruction-level approach and a significant software overhead is expected to
implement fault tolerance at this level. We call a system that uses only this
mechanism therefore a “medium” FCTS.
Fig. 4.3 System recovery times according to the used scheme
34
4 Generalized Algorithm of Fault Tolerance (GAFT)
