Naturally, the more we know about faults and able to recover from them with
less cost the better.
In the used example, the chosen transient fault to permanent fault ratio k is 5,
i.e., transient faults happen five times more often than permanent faults.
Note that even in this conservative case (in practice, k is in the range of 10
3 to
10
5 ) a significant improvement or reliability and overall efficiency of the system can
be shown. A larger k results also in a higher reliability gain.
We did not separate redundancy for checking and redundancy for recovery—it is
subject to further research. Nevertheless, the key outcomes are the following:
– Design of fault-tolerant systems should be revisited: malfunctions must be tolerated at the element level, leaving permanent fault handling to the system level;
– There is an optimum in redundancy level spent on fault tolerance which is less
than 100% and dependent on the anticipated fault coverage level. A duplicated
system is therefore not the most efficient solution;
– Fault coverage defines the efficiency of a concrete fault tolerance
implementation;
– The processes of checking and recovery might be realized concurrently with the
main hardware function, with literally no time delay (time redundancy).
One might argue that the introduced success function does not appropriately
model the real world. A system with, for example, 2% redundancy spent on the
detection and recovery is hardly implementable. In reality, a real success function
might start from say 0.12, where a minimum redundancy level is required to get an
implementable fault-tolerant system.
However, further research is required to define the exact minimum of redundancy, which is needed to implement a system in real life and also to refine the
actual shape of the success function.
The fault coverage is also dependent on the amount of redundancy used in the
system that should be integrated into the analysis.
Another aspect that we did not cover in this analysis is manufacturing. With
adapted manufacturing techniques and conscious placement of the individual
hardware elements such as arithmetic unit, registers, etc. the fault coverage can be
increased and the fault effect on the system lowered.
As there are many open questions, an estimation presented here should be
considered as indicative and not absolute.
4.7 Conceptual Summary
In this chapter, we introduced GAFT, which describes in an algorithmic form a
sequence of required steps to make a system fault tolerant. We then introduced
GAFT in tabular form (template), useful to design a fault-tolerant system and prove
that every step of GAFT is implemented by the chosen design. A specific
44
4 Generalized Algorithm of Fault Tolerance (GAFT)
less cost the better.
In the used example, the chosen transient fault to permanent fault ratio k is 5,
i.e., transient faults happen five times more often than permanent faults.
Note that even in this conservative case (in practice, k is in the range of 10
3 to
10
5 ) a significant improvement or reliability and overall efficiency of the system can
be shown. A larger k results also in a higher reliability gain.
We did not separate redundancy for checking and redundancy for recovery—it is
subject to further research. Nevertheless, the key outcomes are the following:
– Design of fault-tolerant systems should be revisited: malfunctions must be tolerated at the element level, leaving permanent fault handling to the system level;
– There is an optimum in redundancy level spent on fault tolerance which is less
than 100% and dependent on the anticipated fault coverage level. A duplicated
system is therefore not the most efficient solution;
– Fault coverage defines the efficiency of a concrete fault tolerance
implementation;
– The processes of checking and recovery might be realized concurrently with the
main hardware function, with literally no time delay (time redundancy).
One might argue that the introduced success function does not appropriately
model the real world. A system with, for example, 2% redundancy spent on the
detection and recovery is hardly implementable. In reality, a real success function
might start from say 0.12, where a minimum redundancy level is required to get an
implementable fault-tolerant system.
However, further research is required to define the exact minimum of redundancy, which is needed to implement a system in real life and also to refine the
actual shape of the success function.
The fault coverage is also dependent on the amount of redundancy used in the
system that should be integrated into the analysis.
Another aspect that we did not cover in this analysis is manufacturing. With
adapted manufacturing techniques and conscious placement of the individual
hardware elements such as arithmetic unit, registers, etc. the fault coverage can be
increased and the fault effect on the system lowered.
As there are many open questions, an estimation presented here should be
considered as indicative and not absolute.
4.7 Conceptual Summary
In this chapter, we introduced GAFT, which describes in an algorithmic form a
sequence of required steps to make a system fault tolerant. We then introduced
GAFT in tabular form (template), useful to design a fault-tolerant system and prove
that every step of GAFT is implemented by the chosen design. A specific
44
4 Generalized Algorithm of Fault Tolerance (GAFT)
