especially true for software measures [12, 13]. Thus, further analysis of performance/reliability degradation should be taken into account.
The introduced system redundancy might be used in a way to cover only certain
fault types a system can tolerate and thus in terms of fault coverage degrades the
system, but in terms of performance works without difference to the
non-fault-tolerant version.
Software-based redundancy might preserve the same type of fault coverage but
with more time redundancy—delays (recovery time degrades, availability degrades)
[14], but the fault coverage might not, thus the system degrades in terms of
reliability.
3.4 Chapter Conclusion
In this chapter, we introduced the standard reliability theory and showed how to
calculate the reliability of hardware components, assuming Poisson distributed
faults. Based on this, we showed how to calculate the mean time to failure of such
components.
Next, we showed that a component or a system can be made more reliable by
introducing deliberate redundancy for tolerating faults.
We classify redundancy in time, structure, and information, implemented either
in software or hardware. Every concrete redundancy measure, such as duplication,
CRC, etc. can be described using the introduced classification (but not vice versa).
Treating fault tolerance as a process consisting of the basic steps of detection,
location, and elimination of fault, involving hardware and software in cooperation,
proved to be the far superior approach than considering fault tolerance as a static
feature.
We then showed that only by thoroughly analyzing the interconnecting model of
system, model of fault, and model of toleration of this fault using redundancy
becomes possible to derive the “best” fault-tolerant system which fulfills the given
application requirements in a given cost envelope.
References
1. Birolini A (2007) Reliability engineering theory and practice. Springer
2. Von Neumann J (1956) Probabilistic logics and synthesis of reliable organisms from
unreliable components. In: Shannon C, McCarthy J (eds) Automata studies. Princeton
University Press, pp 43–98
3. Pierce WH (1965) Failure-tolerant computer design. Academic Press Inc., New York
4. Laprie J-C (1984) Dependability modeling and evaluation of software and hardware systems.
In: Fehlertolerierende Rechensysteme, 2. GI/NTG/GMR- Fachtagung, London, UK. Springer,
pp 202–215
22
3 Fault Tolerance: Theory and Concepts
Précédent

- 38/315

Suivant