redundancy means such as implementation by hardware redundancy re-execution of
all processor instructions is able to cover multiple steps of GAFT simultaneously.
The template also identifies which redundancy type is used for the implementation of every GAFT step. We then illustrated the use of the template with an
example. Further, we showed that although a complete GAFT implementation
results in a fault-tolerant system, the different solutions vary in different aspects:
performance, reliability, coverage, and cost. The usage scenario of the system
defines the envelope of performance, reliability, and degradation of the system over
time, which must be met by the fault tolerance solution within given cost limits.
Different redundancy schemes have different properties, depending on which
level (instruction, procedure, module, tasks, and system) they are implemented. We
show that as a general guideline, faults should be tolerated as fast and local as
possible (ASAP and ALAP), favoring therefore the instruction level for the
majority of malfunctions.
We extensively covered the three main processes of GAFT. For testing and
checking, we showed that the combination of hardware- and software-based
checking is the most efficient approach:
• Hardware-based checking process works well for covering short time
malfunctions;
• Software-based checking of hardware for latent malfunctions and permanent
errors.
However, for multiple bit and latent transient faults, software must be involved
in addition to the hardware, as a pure hardware solution might be prohibitively
expensive and inefficient. We also conclude that permanent errors should be handled by software as they involve hardware reconfiguration which needs software
support.
A designer of a fault-tolerant system might be tempted to add as much redundancy to the system as possible, thinking that the system is more reliable the more
redundancy is used. With a carefully designed reliability evaluation we were able to
prove that this approach is wrong as the newly introduced redundancy leads to a
decrease in reliability. In fact, we were able to prove that there exists an optimum in
redundancy that should be used for a system, which is in fact less than duplication.
This result has actually a fundamental importance for system design: so far, there
were known approaches in structural and engineering of reliable design and analysis widely applying duplicated, tripled, and quadrupled systems. Mechanical copy
of results from mechanical engineering was wrong—in the electronic design, this
principle is unjustified and makes computer systems absolutely not necessary
redundant and less reliable!
As chapter shows when element of fault-tolerant system is able to tolerate own
malfunctions while system level is dealing with permanent faults of hardware by
reconfiguration the system becomes at the order of magnitude more reliable.
Redundancy of element required to achieve malfunction tolerance at the level of
element was proved to be 12%—much less than duplication of the whole system.
4.7 Conceptual Summary
45
all processor instructions is able to cover multiple steps of GAFT simultaneously.
The template also identifies which redundancy type is used for the implementation of every GAFT step. We then illustrated the use of the template with an
example. Further, we showed that although a complete GAFT implementation
results in a fault-tolerant system, the different solutions vary in different aspects:
performance, reliability, coverage, and cost. The usage scenario of the system
defines the envelope of performance, reliability, and degradation of the system over
time, which must be met by the fault tolerance solution within given cost limits.
Different redundancy schemes have different properties, depending on which
level (instruction, procedure, module, tasks, and system) they are implemented. We
show that as a general guideline, faults should be tolerated as fast and local as
possible (ASAP and ALAP), favoring therefore the instruction level for the
majority of malfunctions.
We extensively covered the three main processes of GAFT. For testing and
checking, we showed that the combination of hardware- and software-based
checking is the most efficient approach:
• Hardware-based checking process works well for covering short time
malfunctions;
• Software-based checking of hardware for latent malfunctions and permanent
errors.
However, for multiple bit and latent transient faults, software must be involved
in addition to the hardware, as a pure hardware solution might be prohibitively
expensive and inefficient. We also conclude that permanent errors should be handled by software as they involve hardware reconfiguration which needs software
support.
A designer of a fault-tolerant system might be tempted to add as much redundancy to the system as possible, thinking that the system is more reliable the more
redundancy is used. With a carefully designed reliability evaluation we were able to
prove that this approach is wrong as the newly introduced redundancy leads to a
decrease in reliability. In fact, we were able to prove that there exists an optimum in
redundancy that should be used for a system, which is in fact less than duplication.
This result has actually a fundamental importance for system design: so far, there
were known approaches in structural and engineering of reliable design and analysis widely applying duplicated, tripled, and quadrupled systems. Mechanical copy
of results from mechanical engineering was wrong—in the electronic design, this
principle is unjustified and makes computer systems absolutely not necessary
redundant and less reliable!
As chapter shows when element of fault-tolerant system is able to tolerate own
malfunctions while system level is dealing with permanent faults of hardware by
reconfiguration the system becomes at the order of magnitude more reliable.
Redundancy of element required to achieve malfunction tolerance at the level of
element was proved to be 12%—much less than duplication of the whole system.
4.7 Conceptual Summary
45
