4.1 The Generalized Algorithm of Fault Tolerance
The traditional process of monitoring faults as an algorithm is generalized using
approaches of [1, 2]. An extension of the algorithm of fault tolerance called the
generalized algorithm of fault tolerance (GAFT) [3–5] is illustrated in Fig. 4.1 and
explained below in more detail.
The primary function of fault monitoring assumes several steps such as follows:
– Detecting faults,
– Identifying faults,
– Identifying faulty component,
– Hardware reconfiguration to achieve a repairable state, and
– Recovery of a correct state(s) for the system and user software.
The process of fault tolerance considers the interaction between HW/SW elements and requires the physics of the fault itself to be presented somehow in the
structure of the algorithm. From the system viewpoint, the nature of hardware faults
is different; a hardware fault is considered as either a permanent fault or a temporary
fault (malfunction).
When fault type determination becomes more complex the complexity of
implementation of processing system grows as well.
This is especially true for multicore systems (MIMD) or processors that support
vector (SIMD) instructions. Pipelined processor implementations make a considerable increase in the implementation complexity also. In principle, an asynchronous hardware fault checking solution can lead to spreading faults within the
whole system and can make recovery therefore unjustifiably complex.
In practice, the complexity of GAFT implementations depends on the complexity of the system, its faults, and the fault tolerance models. Of course, hardware
support for fault tolerance is the fastest way to achieve reliability and with careful
design, there should be little or no performance degradation.
Fig. 4.1 GAFT
26
4 Generalized Algorithm of Fault Tolerance (GAFT)
Précédent

- 41/315

Suivant