The procedure-level scheme requires much less hardware support but, of course,
has a larger timing and software-coding overhead. For the module and the task-level
schemes, extra hardware and system software support for fault tolerance is actually
incomplete as any permanent hardware fault of a crucial HW component would
cause the system to stop working and the GAFT would fail.
The final effort of fault actions to achieve fault tolerance here becomes rebooting
the system to perform the recovery; but again, even with complete system reboot
without hardware redundancy specially dedicated to replace faulty components, the
system can only be resilient to malfunctions (transient faults), not hard faults.
Depending on the implementation level, the systems differ in the required time
frame to achieve fault tolerance (see the thick lines in Fig. 4.3). Note that several
schemes can also be applied in combination, as one scheme alone might not cover
all fault types.
Ideally, all possible levels should be considered for implementation within a
safety-critical system due to large latencies of faults: presence of external impact
inside the system can have a range of up to several seconds (see Sect. 4.2.2). Such a
fault can wrongly be identified as a permanent fault on the instruction level but
correctly handled as a malfunction on higher levels.
Consider the extreme case where the vast majority of malfunctions can be
recovered within the instruction execution and faults that have happened within one
instruction execution are invisible for other instructions (and the software). Then,
only malfunctions with impact time longer than one instruction need to be detected
and recovered by higher levels, for example, at the procedure level. This means that
the vast majority of hardware faults is becoming invisible for the system software
and system software support of hardware deficiency is used rather rarely, which
corresponds to a system according to the left curve in Fig. 4.3.
At the other extreme, the system could tolerate the vast majority of malfunctions
at the task level. The operating system or even user support might be needed to
tolerate hardware faults.
From the user’s point of view, even Microsoft Windows systems might be
considered as fault tolerant as long as an application that was scheduled was
completed and the results are delivered in time, although maintenance hardware or
even a system engineers was involved to fix the fault “in time”. Most Windows
systems assume that system rebooting and restart of the applications are acceptable
ways of operation.
Practical experience in the real-time safety-critical systems domain shows the
opposite as the most critical reliability requirements are an MTTF of 10–25 years
and system availability A(t) of A(t) > 0.99 over the whole system life cycle.
36
4 Generalized Algorithm of Fault Tolerance (GAFT)
Précédent

- 51/315

Suivant