In Sect. 3.1, a short introduction to the redundancy theory related to computer
science is presented. The redundancy theory acts as the basic theoretical guide we
have to follow to get a fault-tolerant system with highest reliability.
However, when designing a fault-tolerant system, other non-computer science
related factors set constraints on the design of such FT systems, for example,
required reliability, which is tightly connected with redundancy of the system and
should be proven using reliability and redundancy theories.
Also, other non-technological factors play a role in the design of FT systems,
especially economics: It is simply not affordable to design and build the system
with highest possible reliability and availability; therefore, trade-offs must be made
in the design and maintenance as this is also a significant cost factor of the overall
system cost.
The central blue box of Fig. 1.1 presents the main principles we follow
throughout our designs and therefore this book:
Reliability. Highest reliability of a safety-critical system is the ultimate goal of
this book and must therefore always be kept in mind.
Simplicity. We believe that the principle of simplicity must be followed in the
design and implementation of SW and HW to achieve the desired reliability and
reconfigurability level of HW and SW. This statement can be easily justified with
the argument that complex systems in themselves are hard to manage and it is
therefore even harder to reach the desired reliability level to make them suitable for
fault-tolerant systems. Consequently, it is important to use only essential redundancy in the system to keep the complexity level low.
Redundancy. Making individual hardware components more reliable is out of
the question for most projects, as the implications also in manufacturing techniques
are severe. We therefore use in this book redundancies of different kinds to improve
the system reliability, a solution which is also in practice often used.
Reconfigurability. Reconfigurability allows the system to adapt twofold: First
to recover from a permanent fault (graceful degradation) and second by adjusting
itself to the requirements of the running application, which helps to keep the system
generic, i.e., suitable for multiple purposes.
Scalability. Keep scalability in mind when designing a system so that it can be
extended, if the requirements change. In this work, scalability does not have highest
priority and might be first sacrificed if one of the other principles is violated.
These principles are well established and especially simplicity and scalability
proved their efficiency through a number of well-known projects and developments
at ETH Zurich [7–15].
These main principles we apply to two different fields: Hardware and System
Software.
In hardware, the above principles are applied at the system level and to the
individual components of the hardware platform. For this, we partition the hardware
in three parts: the Active Zone, the Interface Zone, and the Passive Zone. Each of
these zones has different properties and need therefore different redundancy
mechanism to tolerate faults. We do not include here internal and external devices.
1 Introduction
3
science is presented. The redundancy theory acts as the basic theoretical guide we
have to follow to get a fault-tolerant system with highest reliability.
However, when designing a fault-tolerant system, other non-computer science
related factors set constraints on the design of such FT systems, for example,
required reliability, which is tightly connected with redundancy of the system and
should be proven using reliability and redundancy theories.
Also, other non-technological factors play a role in the design of FT systems,
especially economics: It is simply not affordable to design and build the system
with highest possible reliability and availability; therefore, trade-offs must be made
in the design and maintenance as this is also a significant cost factor of the overall
system cost.
The central blue box of Fig. 1.1 presents the main principles we follow
throughout our designs and therefore this book:
Reliability. Highest reliability of a safety-critical system is the ultimate goal of
this book and must therefore always be kept in mind.
Simplicity. We believe that the principle of simplicity must be followed in the
design and implementation of SW and HW to achieve the desired reliability and
reconfigurability level of HW and SW. This statement can be easily justified with
the argument that complex systems in themselves are hard to manage and it is
therefore even harder to reach the desired reliability level to make them suitable for
fault-tolerant systems. Consequently, it is important to use only essential redundancy in the system to keep the complexity level low.
Redundancy. Making individual hardware components more reliable is out of
the question for most projects, as the implications also in manufacturing techniques
are severe. We therefore use in this book redundancies of different kinds to improve
the system reliability, a solution which is also in practice often used.
Reconfigurability. Reconfigurability allows the system to adapt twofold: First
to recover from a permanent fault (graceful degradation) and second by adjusting
itself to the requirements of the running application, which helps to keep the system
generic, i.e., suitable for multiple purposes.
Scalability. Keep scalability in mind when designing a system so that it can be
extended, if the requirements change. In this work, scalability does not have highest
priority and might be first sacrificed if one of the other principles is violated.
These principles are well established and especially simplicity and scalability
proved their efficiency through a number of well-known projects and developments
at ETH Zurich [7–15].
These main principles we apply to two different fields: Hardware and System
Software.
In hardware, the above principles are applied at the system level and to the
individual components of the hardware platform. For this, we partition the hardware
in three parts: the Active Zone, the Interface Zone, and the Passive Zone. Each of
these zones has different properties and need therefore different redundancy
mechanism to tolerate faults. We do not include here internal and external devices.
1 Introduction
3
