If the system can recover itself and continue working in a degraded mode after
the occurrence of a fault, it is called a Gracefully Degradable System (GDS) [8]. In
turn, if a system can stop itself correctly in an acceptable state once a fault has been
detected, it is called a Fail-Stop System (FSS) [9].
The definition of fault-tolerant systems proposed for RT systems in [1] is more
rigorous and uses an algorithmic definition of the feature: A system is called fault
tolerant if and only if it implements GAFT.
In other words:
An RT system is called fault tolerant if it provides the full application functionality,
executes GAFT (including recovery) when necessary within a clearly specified
timeframe (between sequential data transmissions or tasks) transparently for the
applications.
Applied to the example of an airplane flight system, this would mean that if the
system can recover itself within defined time (between sequential data inputs, in
practice about 125 ms) in failure-free mode (or with recovery that is transparent for
the application) then this system is called fault tolerant.
4.3 Example of Possible GAFT Implementation
For every step in GAFT, at least one redundancy-handling mechanism is required to
make the system fault tolerant. In this chapter, we give an example of a possible
concept that illustrates the process. Note that one redundancy scheme can cover
more than one step in GAFT, as one redundancy feature might provide means to
detect faults, but also treat faults. Triplicate memory would be one example.
Table 4.2 shows the complete GAFT with used redundancy types to implement
each algorithm step for our example application of a flight control system. The
numbers in the redundancy columns correspond to the redundancy types applied in
each step. These numbers, in turn, are briefly described in Table 4.3 to illustrate the
functionality.
Of course, not all fault types can be covered by each of these redundancy
mechanisms. We concentrate here on the HW faults presented in Sect. 4.3.4. Refer
to [8] for a more detailed classification of possible faults not only in HW but also in
SW and during the design phase.
Table 4.2 shows that the various applied redundancy mechanisms are used in
single or across several steps of the GAFT algorithm. The numbers in the redundancy-type fields correspond to the redundancy mechanisms presented in Table 4.3.
All mechanisms combined result in a system that implements all steps of GAFT and
is, therefore, fault tolerant.
Not all components, however, are in this example equally strong fault tolerant.
For example, the processor and the memory subsystem are massively covered by
checking schemes, but other systems such I/O devices are only covered from faults
4.2 Definition of Fault Tolerance by GAFT
31
the occurrence of a fault, it is called a Gracefully Degradable System (GDS) [8]. In
turn, if a system can stop itself correctly in an acceptable state once a fault has been
detected, it is called a Fail-Stop System (FSS) [9].
The definition of fault-tolerant systems proposed for RT systems in [1] is more
rigorous and uses an algorithmic definition of the feature: A system is called fault
tolerant if and only if it implements GAFT.
In other words:
An RT system is called fault tolerant if it provides the full application functionality,
executes GAFT (including recovery) when necessary within a clearly specified
timeframe (between sequential data transmissions or tasks) transparently for the
applications.
Applied to the example of an airplane flight system, this would mean that if the
system can recover itself within defined time (between sequential data inputs, in
practice about 125 ms) in failure-free mode (or with recovery that is transparent for
the application) then this system is called fault tolerant.
4.3 Example of Possible GAFT Implementation
For every step in GAFT, at least one redundancy-handling mechanism is required to
make the system fault tolerant. In this chapter, we give an example of a possible
concept that illustrates the process. Note that one redundancy scheme can cover
more than one step in GAFT, as one redundancy feature might provide means to
detect faults, but also treat faults. Triplicate memory would be one example.
Table 4.2 shows the complete GAFT with used redundancy types to implement
each algorithm step for our example application of a flight control system. The
numbers in the redundancy columns correspond to the redundancy types applied in
each step. These numbers, in turn, are briefly described in Table 4.3 to illustrate the
functionality.
Of course, not all fault types can be covered by each of these redundancy
mechanisms. We concentrate here on the HW faults presented in Sect. 4.3.4. Refer
to [8] for a more detailed classification of possible faults not only in HW but also in
SW and during the design phase.
Table 4.2 shows that the various applied redundancy mechanisms are used in
single or across several steps of the GAFT algorithm. The numbers in the redundancy-type fields correspond to the redundancy mechanisms presented in Table 4.3.
All mechanisms combined result in a system that implements all steps of GAFT and
is, therefore, fault tolerant.
Not all components, however, are in this example equally strong fault tolerant.
For example, the processor and the memory subsystem are massively covered by
checking schemes, but other systems such I/O devices are only covered from faults
4.2 Definition of Fault Tolerance by GAFT
31
