The introduction of static redundancy in hardware and system software (HW/
SSW) might be prohibitively expensive; therefore, it is much better to introduce a
process that implements fault tolerance assuming dynamic interaction of existing
redundancy types between elements, as it is illustrated in Fig. 3.4.
The main components of the system (hardware and system software) are
themselves sources of possible internal hardware faults. At the same time, using
various redundancy types, fault tolerance of the computer system can be achieved
both by design and by operation. Thus, the elements of the problem become the
elements of the solution.
For the implementation of fault tolerance as a process, literature usually considers a three-step fault-handling algorithm: detection—location—reconfiguration
[11] which was extended and broadened to a more general scheme. This algorithm
should be applied for each part of the system, i.e., system software and hardware.
The implementation of the algorithm requires the use of the above introduced three
different redundancy types (Fig. 3.4) and will be further discussed with necessary
details.
In fact, a fault tolerance can be considered as three basic blocks that determine a
new feature: the system model, the model of faults that need to be tolerated, and the
fault tolerance model. This includes both processes: the design and the implementation of the FT system. In the next section, the fault tolerance process models
are described in more detail.
3.3 Models for Fault Tolerance
Say M is the known model of the system to perform a given function F. To this
model, we introduce a new feature that was not defined before: extreme reliability.
To express the existence of reliability in the system, the predicates P and Q are
introduced to determine the state of the model with regard to the new quality. P and
Q also define the direction of the time arrow (see Fig. 3.5).
To analyze ways how to achieve a required reliability level with performance
and power consumption constraints, we offer a combination of the following three
models:
– The model of an object M o or M system, in this case, the computer system.
– The model of the faults M fault that an RT FT system should tolerate.
– The model (scheme) M FT or new structure that implements fault tolerance.
The system model, fault model, and fault tolerance model are mutually dependent
as it is shown at the bottom of Fig. 3.5. Note that in the here presented approach,
the development and manufacturing cost of a solution is not considered.
M fault is a description of all possible faults a system must tolerate. In binary
logic, a typical permanent fault manifests as “stuck at zero” or “stuck at one”.
Behavioral faults such as Byzantine faults (malfunctions) and hidden faults
(so-called latent faults) that exist in the hardware over a long period of time do not
18
3 Fault Tolerance: Theory and Concepts
SSW) might be prohibitively expensive; therefore, it is much better to introduce a
process that implements fault tolerance assuming dynamic interaction of existing
redundancy types between elements, as it is illustrated in Fig. 3.4.
The main components of the system (hardware and system software) are
themselves sources of possible internal hardware faults. At the same time, using
various redundancy types, fault tolerance of the computer system can be achieved
both by design and by operation. Thus, the elements of the problem become the
elements of the solution.
For the implementation of fault tolerance as a process, literature usually considers a three-step fault-handling algorithm: detection—location—reconfiguration
[11] which was extended and broadened to a more general scheme. This algorithm
should be applied for each part of the system, i.e., system software and hardware.
The implementation of the algorithm requires the use of the above introduced three
different redundancy types (Fig. 3.4) and will be further discussed with necessary
details.
In fact, a fault tolerance can be considered as three basic blocks that determine a
new feature: the system model, the model of faults that need to be tolerated, and the
fault tolerance model. This includes both processes: the design and the implementation of the FT system. In the next section, the fault tolerance process models
are described in more detail.
3.3 Models for Fault Tolerance
Say M is the known model of the system to perform a given function F. To this
model, we introduce a new feature that was not defined before: extreme reliability.
To express the existence of reliability in the system, the predicates P and Q are
introduced to determine the state of the model with regard to the new quality. P and
Q also define the direction of the time arrow (see Fig. 3.5).
To analyze ways how to achieve a required reliability level with performance
and power consumption constraints, we offer a combination of the following three
models:
– The model of an object M o or M system, in this case, the computer system.
– The model of the faults M fault that an RT FT system should tolerate.
– The model (scheme) M FT or new structure that implements fault tolerance.
The system model, fault model, and fault tolerance model are mutually dependent
as it is shown at the bottom of Fig. 3.5. Note that in the here presented approach,
the development and manufacturing cost of a solution is not considered.
M fault is a description of all possible faults a system must tolerate. In binary
logic, a typical permanent fault manifests as “stuck at zero” or “stuck at one”.
Behavioral faults such as Byzantine faults (malfunctions) and hidden faults
(so-called latent faults) that exist in the hardware over a long period of time do not
18
3 Fault Tolerance: Theory and Concepts
