Active Zone. The active zone consists of the arithmetic unit and the logic unit,
both separated for better fault isolation and easier implementation of hardware tests.
Interface Zone. This zone includes all communication components, such as the
processor bus, memory bus, etc. The interface zone must be protected from external
radiation and needs therefore mechanisms to detect and recover from faults.
A configurable bus allows the reconfiguration of the hardware to exclude failed
hardware components and go into a degraded state, or replaces the failed component with a working one.
Passive Zone. These include basically the storage systems, such as memory that
do not act by themselves but are used by controllers or other data processing
devices. Here again, the hardware should be able to detect faults and recover from
them, but as they are passive systems, the actual fault tolerance must be implemented in conjunction with the interface zone, i.e., the bus.
In software, we distinguish the following parts:
Semantic. The Graph Logic Model (GLM) [16] and Graph Logic Language
(GLL) [17] are an interesting attempt for the description of system dependencies,
and thus provide a framework for the design and development of the required fault
tolerance in HW and SSW. The GLL and GLM are however still in their early
phases, and are at the time of writing still subject for further research.
Structure. The programming language is of importance to support safety-critical
features. Especially code safety can benefit from a strong types of language without
direct assembler support, to force the programmer to write safe code. Another
aspect here is the support for recovery points as recovery point creation can directly
exploit language features (memory space partitioning, etc.).
Runtime. The main part of this work covers runtime system aspects. In addition
to standard runtime system tasks such as resource management, scheduling, and
inter-process communication, a runtime system supporting fault tolerance must
cover additional tasks such as hardware and software state monitoring to act upon
failing components or detected faults. Especially to support recovery from faults,
additional features such as recovery points, recovery itself, and recovery monitoring
are needed.
The modern requirements for hardware and system software for FT RT systems
are as follows:
In hardware:
– High Performance—32 bit, 1–4Ghz,
– Highest possible reliability and availability,
– Means to detect and handle faults and errors (fault tolerance),
– “Zero” maintenance over the operational life span (for satellites, aircrafts, etc.),
– Intelligent power design with backup battery for fault tolerance during
operation,
– Mechanical and vibration resistance,
– Graceful mechanical degradation, and
– Feasibility.
4
1 Introduction
both separated for better fault isolation and easier implementation of hardware tests.
Interface Zone. This zone includes all communication components, such as the
processor bus, memory bus, etc. The interface zone must be protected from external
radiation and needs therefore mechanisms to detect and recover from faults.
A configurable bus allows the reconfiguration of the hardware to exclude failed
hardware components and go into a degraded state, or replaces the failed component with a working one.
Passive Zone. These include basically the storage systems, such as memory that
do not act by themselves but are used by controllers or other data processing
devices. Here again, the hardware should be able to detect faults and recover from
them, but as they are passive systems, the actual fault tolerance must be implemented in conjunction with the interface zone, i.e., the bus.
In software, we distinguish the following parts:
Semantic. The Graph Logic Model (GLM) [16] and Graph Logic Language
(GLL) [17] are an interesting attempt for the description of system dependencies,
and thus provide a framework for the design and development of the required fault
tolerance in HW and SSW. The GLL and GLM are however still in their early
phases, and are at the time of writing still subject for further research.
Structure. The programming language is of importance to support safety-critical
features. Especially code safety can benefit from a strong types of language without
direct assembler support, to force the programmer to write safe code. Another
aspect here is the support for recovery points as recovery point creation can directly
exploit language features (memory space partitioning, etc.).
Runtime. The main part of this work covers runtime system aspects. In addition
to standard runtime system tasks such as resource management, scheduling, and
inter-process communication, a runtime system supporting fault tolerance must
cover additional tasks such as hardware and software state monitoring to act upon
failing components or detected faults. Especially to support recovery from faults,
additional features such as recovery points, recovery itself, and recovery monitoring
are needed.
The modern requirements for hardware and system software for FT RT systems
are as follows:
In hardware:
– High Performance—32 bit, 1–4Ghz,
– Highest possible reliability and availability,
– Means to detect and handle faults and errors (fault tolerance),
– “Zero” maintenance over the operational life span (for satellites, aircrafts, etc.),
– Intelligent power design with backup battery for fault tolerance during
operation,
– Mechanical and vibration resistance,
– Graceful mechanical degradation, and
– Feasibility.
4
1 Introduction
