In our case, input validation is used to detect faulty data, which is rejected to
prevent the faults from entering the system. If such an error is detected, the self-test
procedure of the input data hardware device has to run to diagnose the fault type
(transient or permanent).
The numbers used in the table are explained in Table 4.3. Table 4.3 shows all
used redundancy mechanisms in our example above with a short description. Not
every aspect of such a flight control system can be covered by GAFT, for example,
real-time aspects are not present.
Thus, GAFT covers only the redundancy part of a possible implementation and
other techniques, such as scheduling strategies are used to guarantee real-time
constraints.
The combination of the two features: fault tolerance and real-time in one system
is very challenging as they work against each other: real time requires tasks to have
finished processing before their individual deadlines; in turn, fault tolerance,
especially implemented using redundancies HW(T) and SW(T), requires additional
time to complete GAFT.
Fault tolerance implementation using time redundancy therefore is hardly possible for real-time systems. For extremely time-critical tasks, application of HW(T)
and SW(T) for fault tolerance should therefore be omitted.
The combination of duplicated storage devices (HW(2S)) and CRC-32 (SW(i))
is a typical example of combining two redundancies which on their own can only
detect faults of a storage subsystem. The duplicated storage is used to detect faults,
whereas the SW-based CRC-32 is used to identify the correct data, which is then
used to correct the faulty instance.
4.4 GAFT Properties: Performance, Reliability, Coverage
As already mentioned above and shown in Fig. 4.1 and Table 4.1, the three connected processes checking and testing, preparation for recovery and recovery might
be implemented by and within SSW and HW at the design phase and runtime phase
of the whole system life cycle.
Obviously, different implementations of the three processes differ in terms of
fault coverage, achievable reliability, availability, and cost. Different implementations of GAFT vary in terms of used redundancy types. This, therefore, changes the
time to complete GAFT.
The runtime phase T of the system life cycle in terms of time redundancy fault
tolerance can be considered at different levels of granularity that are related to the
scope of the program being executed. We differentiate the following five levels:
instruction, procedure, module, task, and system (not shown) as shown in Fig. 4.3.
The instruction-level scheme assumes that when a fault appears, its influence is
eliminated within the instruction execution, using hardware redundancy for fault
detection, fault location, and fault recovery.
4.3 Example of Possible GAFT Implementation
33
Précédent

- 48/315

Suivant