externally, or by a signal from an element, which is part of the model, indicating
that there is a discrepancy in behavior of one or more elements.
5.6 Chapter Summary
In this chapter, we presented the model of active system safety, which applied to
fault-tolerant systems serves as an extension to GAFT, i.e., it shows how to deal
with latent faults that stay for a long time in the system until they manifest
somewhere. To identify the true source of a fault, we propose therefore to map all
elements of the system and its dependencies to a dependency matrix.
Faults can only be spread in the system along element dependencies; therefore,
the shown tracing algorithms can trace a fault manifestation (error) to the true
source of the fault. The appropriate recovery actions are stored in the recovery
matrix, which is then executed to recover the system, i.e., eliminate the fault. In
reverse direction, we showed that the tracing algorithm can also be used to derive
the set of elements that might have been affected by the fault and run diagnostic
routines to ensure the integrity of the elements.
Conclusion for Part I, Chaps. 1–5
– In Part I, we introduced the topic of hardware faults and showed that
safety-critical systems, especially in high altitude, are prone to faults by external
events and that the number of these elements is expected to rise in the future.
Based on this finding, we using the standard reliability theory show how reliability and fault tolerance are connected, i.e., the deliberate introduction of
redundancy increases reliability if the reliability benefit resulting from the redundancy scheme use exceeds the decrease in reliability due to the actual
implementation of the redundancy scheme itself.
– We then presented our fault tolerance model, where we showed that the three
models: system, fault, and fault tolerance are connected and only the thorough
analysis of the system and the expected faults lead to the efficient fault tolerance
solution. GAFT serves as a model and guide for the implementation of fault
tolerance and defines all required steps to make a system fault tolerant.
– GAFT in a tabular form which serves as a template for the implementation of a
fault tolerance solution. We defined that a system is fault tolerant when it
implements every step of GAFT. We will further go into details about possible
implementations in Part II of this work. Different fault tolerance solutions have
different properties in performance, reliability, fault coverage, and cost and are
constraint by the requirements of the usages scenario of the system.
– In our reliability evaluation of hardware redundancy, we clearly showed that
there exists a optimum in achievable reliability gain using redundancy which is
in fact less than duplication. Therefore, the usual approach to flight control
systems in airplanes (Airbus, Boeing), i.e., massive redundancy using triplicated
or even quadruplicated systems does not result in the most reliable solution.
56
5 GAFT Generalization: A Principle and Model of Active System…
Précédent

- 71/315

Suivant