Chapter 3
Fault Tolerance: Theory and Concepts
Abstract This chapter briefly introduces how reliability of the system might be
considered in combination with fault tolerance. Having introduced hardware faults
in the previous chapter, we present in this chapter the elements of theory of fault
tolerance and reliability and show how the hardware components of a computing
system can be made more resilient to hardware faults. We then introduce the
mathematical definition of reliability and show how to calculate the reliability of a
system according to the topology of its components. Then we describe the connection between reliability and fault tolerance, i.e., we show how applying different
types of redundancy, implemented in software and hardware, increases the reliability of a system. Also, some design advices are given.
3.1 Introduction to Reliability Theory
In addition to the just shown problems of faults, the HW system design for RT FT
applications becomes more difficult due to the limited power envelope and
requirements such as extremely reliable RT data storage, survivable packaging,
possible data retrieval at any time, and a variable number of external influences
such as radiation or vibration.
This results in difficult challenges in terms of power dissipation, heat management, mechanical resilience, and of course malfunction tolerance. Given the
requirement of zero maintenance for onboard computer hardware and system
software, the design and reliability of the FT system is extremely challenging.
There are two different known theoretical approaches to meet the abovementioned reliability requirements and specifications:
1. By developing extremely reliable devices with a substantially higher Mean Time
To Failure (MTTF) and substantially higher than expected lifetime of the
monitored object. In case of aviation application, an airplane control flight
system should have a substantially higher availability and lifetime than the
airplane itself. Birolini [1] introduced a comprehensive theoretical approach
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_3
11
Fault Tolerance: Theory and Concepts
Abstract This chapter briefly introduces how reliability of the system might be
considered in combination with fault tolerance. Having introduced hardware faults
in the previous chapter, we present in this chapter the elements of theory of fault
tolerance and reliability and show how the hardware components of a computing
system can be made more resilient to hardware faults. We then introduce the
mathematical definition of reliability and show how to calculate the reliability of a
system according to the topology of its components. Then we describe the connection between reliability and fault tolerance, i.e., we show how applying different
types of redundancy, implemented in software and hardware, increases the reliability of a system. Also, some design advices are given.
3.1 Introduction to Reliability Theory
In addition to the just shown problems of faults, the HW system design for RT FT
applications becomes more difficult due to the limited power envelope and
requirements such as extremely reliable RT data storage, survivable packaging,
possible data retrieval at any time, and a variable number of external influences
such as radiation or vibration.
This results in difficult challenges in terms of power dissipation, heat management, mechanical resilience, and of course malfunction tolerance. Given the
requirement of zero maintenance for onboard computer hardware and system
software, the design and reliability of the FT system is extremely challenging.
There are two different known theoretical approaches to meet the abovementioned reliability requirements and specifications:
1. By developing extremely reliable devices with a substantially higher Mean Time
To Failure (MTTF) and substantially higher than expected lifetime of the
monitored object. In case of aviation application, an airplane control flight
system should have a substantially higher availability and lifetime than the
airplane itself. Birolini [1] introduced a comprehensive theoretical approach
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_3
11
