4.6 Hardware Redundancy and Reliability
Consider a system with a given reliability which is prone to transient and permanent
faults (Fig. 4.5, left graph). If additional hardware is added to detect transient faults,
the introduced redundancy to achieve fault detection reduces the absolute reliability
of the system as more hardware is used that is prone to faults.
At the same time, if the introduced redundancy is not only used to detect
transient faults but also to tolerate them, the reliability of the system increases
(Fig. 4.5, right graph). Thus, part of the problem—the decreasing reliability of
hardware caused by redundancy at some point (after introduction of recoverability
process P3) becomes part of a solution. Note that the system is however still prone
to permanent faults.
Figure 4.5 illustrates reliability degradation in presence of malfunction and
permanent faults and gain for the system assuming permanent fault ration k = 10
−5 ,
coefficient of malfunction to permanent fault ratio k = {0, 1000}, and redundancy
of hardware d = 0.12. The data applied are natural based on our own design of
processor with malfunction tolerance [13], while ratio introduced deliberately lower
than real, in real practice as ratio of malfunction to permanent faults, varies from
10
5 to 10
7 .
Clear that analysis of effect of redundancy on the reliability of a system is worth
to clarify a bit more. Note that malfunction tolerance as mentioned before can be
Fig. 4.4 Trade-offs to be made in fault-tolerant system design: time-, performance-, and
reliability-wise
38
4 Generalized Algorithm of Fault Tolerance (GAFT)
Consider a system with a given reliability which is prone to transient and permanent
faults (Fig. 4.5, left graph). If additional hardware is added to detect transient faults,
the introduced redundancy to achieve fault detection reduces the absolute reliability
of the system as more hardware is used that is prone to faults.
At the same time, if the introduced redundancy is not only used to detect
transient faults but also to tolerate them, the reliability of the system increases
(Fig. 4.5, right graph). Thus, part of the problem—the decreasing reliability of
hardware caused by redundancy at some point (after introduction of recoverability
process P3) becomes part of a solution. Note that the system is however still prone
to permanent faults.
Figure 4.5 illustrates reliability degradation in presence of malfunction and
permanent faults and gain for the system assuming permanent fault ration k = 10
−5 ,
coefficient of malfunction to permanent fault ratio k = {0, 1000}, and redundancy
of hardware d = 0.12. The data applied are natural based on our own design of
processor with malfunction tolerance [13], while ratio introduced deliberately lower
than real, in real practice as ratio of malfunction to permanent faults, varies from
10
5 to 10
7 .
Clear that analysis of effect of redundancy on the reliability of a system is worth
to clarify a bit more. Note that malfunction tolerance as mentioned before can be
Fig. 4.4 Trade-offs to be made in fault-tolerant system design: time-, performance-, and
reliability-wise
38
4 Generalized Algorithm of Fault Tolerance (GAFT)
