In turn, a combination of HW and SSW implementation should be quite effective
depending on the variety of tolerated faults and the acceptable level of performance
degradation. It is clear that GAFT should deal with malfunctions first because they
are easier to recover than permanent faults.
The closer the fault type detection is to the start of the algorithm of fault tolerance, the faster the algorithm completes in the case of malfunctions. Taking into
account that the ratio of malfunctions to permanent faults is roughly 10
4
–10
5 a
differentiation based on fault types will obviously provide an improvement of FT
efficiency. The algorithm of fault tolerance is shown in Fig. 4.1.
GAFT is initiated by external events, such as fault detection, periodic testing, or
maintenance. Fault detection can be performed synchronously or asynchronously
using hardware or software redundancy. Further, this chapter discusses GAFT
implementation and activation by hardware and system software in more detail.
An example of asynchronous hardware redundancy is HW(∂I)—redundant data
bits to check data errors. In the first step, the errors are detected and in a second
step, the redundancy information is used to recover the system.
An example of synchronous hardware redundancy would be HW(∂T)—special
hardware delay schemes to stabilize signals and avoid malfunctions.
Other redundancy examples may also be used for the same purpose including,
for example, SW (∂S, ∂T)—execution of SW-based test programs to validate
hardware integrity performed during idle time (then SW(∂T) = 0 from the system
point of view) or SW(∂I)—informational redundancy of the program.
Hardware faults detected by hardware or software invoke specific
fault-dependent actions and initiate the execution of a diagnosing program to
identify the fault type and to locate a specific faulty chip. If the fault is permanent,
the component that contains the fault is excluded from the system by a special step
of GAFT called hardware reconfiguration.
As long as the hardware design supports this and there are still redundant
resources available, the system is reconfigured into a properly working state. If no
redundant resources are available, the faulty component can still be excluded from
the system as long as it is not essential, and the system continues processing in a
degraded state.
Hardware fault tolerance from permanent faults requires redundant hardware
components HW(S 1 , S 1 ) and is unavoidable for FT RT systems. Hardware reconfiguration must also be performed synchronously and invisible, if possible, to the
system function to support continuous operation. If, due to already performed
reconfiguration, it is no longer possible to further reconfigure the system, the faulty
component might also be disabled, and the processing could continue working as
long as no crucial component such as the processor or the memory is affected.
Analog to the hardware recovery, the software must also be recovered, i.e.,
proving that the fault did not affect the software state or excluding the effects of the
fault if it had an impact. It is also necessary to make the system aware of a new
possible hardware topology, in the case that an HW component failed and could not
be recovered.
4.1 The Generalized Algorithm of Fault Tolerance
27
Précédent

- 42/315

Suivant