For the module level of recovery, a fault and its influence is tolerated within a
module execution. For example, module restart or run of a simplified alternative
module might be used to avoid system restart, reboot, etc. The same comments
apply as for the procedure case, except that time and program overheads for
achieving fault tolerance are even higher as the state space is likely to be much
larger.
Redundancy on the module level might be considered as SW(S, I, T) where S
stands for extra software element to perform hardware checking (if hardware
schemes are unavailable) and prepare recovery point before module run. Other
redundancies I and T define the required extra information to form a recovery point
for the module and time required to generate it.
It was already shown in the late 80s that the process of checking and recovery
can be implemented in parallel [3, 4], thus T ! 0 from the system point of view, so
it could also be treated as a supportive redundancy.
The task-level scheme eliminates a hardware fault and its influence by a task
restart after the hardware reconfiguration. Here, redundancy is obviously required
much greater than in previously described schemes.
Finally, the system state may be recovered by a reboot and task repetition. Due
to the high recovery time, we call such a system a “weak” FCTS.
Recovery on the level of the system is implicitly available on all systems, as
rebooting the system in case of an error basically performs recovery. This corresponds to turning on a system on and must, therefore, be supported in any case. We
consider this as the “last resort” if all other measures fail.
The initialization of these GAFT scheme implementations might also require
some time especially on the higher recovery levels or depends on hardware signals
from checking schemes.
The time impact of implementing GAFT at each of these levels is different, as
well as delays caused by their use. Current embedded system practice indicates the
following orders of magnitudes for timing:
– Nanoseconds for the instruction level,
– Microseconds to milliseconds for the procedure level,
– Hundreds of milliseconds to seconds for the module level, and
– Seconds to tens of seconds at the task level.
The different schemes have different overheads, capabilities for tolerating various fault classes, power consumption overheads, and system costs. For the
instruction-level scheme, the required hardware support (such as duplicate or
triplicate hardware modules) could result in serious power consumption overheads,
size, and cost.
However, not all hardware subsystems need necessarily be designed in this way.
For example, a cost–benefit analysis (where “cost” infers financial, power consumption, chip area, and time delays) might indicate that it is worth having the
processor and system RAM protected at this level, but not the other subsystems.
4.4 GAFT Properties: Performance, Reliability, Coverage
35
module execution. For example, module restart or run of a simplified alternative
module might be used to avoid system restart, reboot, etc. The same comments
apply as for the procedure case, except that time and program overheads for
achieving fault tolerance are even higher as the state space is likely to be much
larger.
Redundancy on the module level might be considered as SW(S, I, T) where S
stands for extra software element to perform hardware checking (if hardware
schemes are unavailable) and prepare recovery point before module run. Other
redundancies I and T define the required extra information to form a recovery point
for the module and time required to generate it.
It was already shown in the late 80s that the process of checking and recovery
can be implemented in parallel [3, 4], thus T ! 0 from the system point of view, so
it could also be treated as a supportive redundancy.
The task-level scheme eliminates a hardware fault and its influence by a task
restart after the hardware reconfiguration. Here, redundancy is obviously required
much greater than in previously described schemes.
Finally, the system state may be recovered by a reboot and task repetition. Due
to the high recovery time, we call such a system a “weak” FCTS.
Recovery on the level of the system is implicitly available on all systems, as
rebooting the system in case of an error basically performs recovery. This corresponds to turning on a system on and must, therefore, be supported in any case. We
consider this as the “last resort” if all other measures fail.
The initialization of these GAFT scheme implementations might also require
some time especially on the higher recovery levels or depends on hardware signals
from checking schemes.
The time impact of implementing GAFT at each of these levels is different, as
well as delays caused by their use. Current embedded system practice indicates the
following orders of magnitudes for timing:
– Nanoseconds for the instruction level,
– Microseconds to milliseconds for the procedure level,
– Hundreds of milliseconds to seconds for the module level, and
– Seconds to tens of seconds at the task level.
The different schemes have different overheads, capabilities for tolerating various fault classes, power consumption overheads, and system costs. For the
instruction-level scheme, the required hardware support (such as duplicate or
triplicate hardware modules) could result in serious power consumption overheads,
size, and cost.
However, not all hardware subsystems need necessarily be designed in this way.
For example, a cost–benefit analysis (where “cost” infers financial, power consumption, chip area, and time delays) might indicate that it is worth having the
processor and system RAM protected at this level, but not the other subsystems.
4.4 GAFT Properties: Performance, Reliability, Coverage
35
