Reliability is concerned with the duration of normal life of an ordinary nonredundant unit, whereas fault tolerance enables the possibility of “life after death”. It
is assumed that there is a mechanism available that can just in time, when the fault
is detected, repair the unit immediately and completely. Furthermore, the system
repair is fast enough that an external observer of the operational system doesn’t
even see that the fault has occurred and has been repaired.
In other words, fault tolerance assumes that fault manifestation; fault identification, if necessary; reconfiguration; and recovery are transparent to the system.
After the appearance of a permanent fault, a fault-tolerant system will be dynamically restored and thought “as good as new” in operational terms, except for the
fact that some of the redundancy has been used up and this may limit the possibilities for future repairs.
The classic parallel generalization of the standard redundancy model [1]
describes a system of n statistically identical elements in active redundancy, when k
elements are required to perform the system function and the remaining n − k
elements are in reserve. The function of the system is considered as successful if
during a certain time frame at least k elements of the system were available.
For a simpler example such as a 1 out of 2 system shown in Fig. 3.2, the system
function is complete if at least one of the elements is known to be working correctly. The second element is redundant and only introduced for reliability purposes
when the first unit is known to be faulty.
For the system in Fig. 3.2, the following reliability function can be derived:
R s ¼ R 1 ðtÞ þ R 2 ðtÞ À R 1 ðtÞR 2 ðtÞ
ð 3:7Þ
Thus, it’s obvious that redundancy even for this classic case could improve the
reliability of the system considerably. Note that the redundant elements need not
necessarily correspond to identical elements, but could also correspond to additional hardware used to detect and treat malfunctions.
The hardware fault tolerance and performance, as already mentioned, are limited
by available technologies and the implemented hardware and system software
solutions. This results in a system design process with constraints in terms of
applied approaches to the above-discussed case with n elements, whereas the
maximum reliability is achieved if all elements have equal reliability (Fig. 3.3).
Therefore, the new functionality of the system software which is required to
make a system fault-tolerant and real-time capable is, above all, the support for
equal reliability of the main components of the hardware architecture (left box of
Fig. 3.2 Parallel system
structure
14
3 Fault Tolerance: Theory and Concepts
Précédent

- 30/315

Suivant