inability to repeat the same impact, for example, of alpha particles and the need to
use a special purposes build radiation facilities.
To test the software, the use of a simulator often proves to be far cheaper and
much more efficient as specific faults can be modeled in software. In this case, real
radiation tests are only used to test the hardware and to verify the chosen fault
models.
The second life cycle phase, namely, the operational phase of a safety-critical
system is in principal similar to a standard embedded system with one big difference, where maintenance of a standard embedded system in case of failure is
usually possible although costly, and the situation typically changes for
safety-critical systems.
If one of these systems fails, the consequences can be the loss of human lives
(failure of a flight control system in an airplane) or physical maintenance might not
even be possible (satellite system). Therefore, maintenance, especially physical
maintenance cannot be considered as a part of life cycle for this kind of systems.
The maintenance interval of a system is usually derived from the life span of
such a system, in other words, from the MTTF of such a system, which can often
not clearly be derived. An optimal solution would be a system that has the power of
self-diagnosis up to the level where it can predict its failure in the near future. In this
case, it could inform the operator which in turn could initiate the maintenance
procedure before the system fails. This would lower maintenance cost significantly.
6.2 System Software Phases
System software consists of the operating system and all other hardware-related
software, which provides services to user programs. It is thus responsible for
abstracting and managing hardware resources and gives application software controlled access to these resources. Some examples are interruption handlers, memory
management, device drivers, scheduler, garbage collector, etc.
These traditional system software components are already extensively explained
in literature [10–12]; therefore, we do not go into too much details about these
functionalities, but introduce here instead some new system software concepts
which are required in the field of fault-tolerant computing.
As already shown, hardware is suspect to malfunctions and failures despite all
efforts on the hardware side. Responsible of managing faulty hardware should be
the operating system itself as application programs have no direct access to the
hardware and usually lack information to deal with hardware errors. Thus, the
hardware state (see Chap. 7) must be represented in the system software or even in
the programming language. We state here that a fault-tolerant system must,
according to GAFT, support the following three additional processes:
Checking. The system software must collaborate with the hardware to perform the
system checking, i.e., show that the hardware is fault free. Checking procedures can
6.1 System Software Life Cycle Versus Fault Tolerance
67
Précédent

- 81/315

Suivant