subject of special research at a later date, taking into account complexity of the task
dependency organization.
6.1 System Software Life Cycle Versus Fault Tolerance
The system life cycle of an FT system can be split into two phases: first, the
development phase (design and implementation) and second the operational phase
(deployment, operating use, and maintenance).
The first, and in terms of an FT system almost most important part is the design
phase. In this phase, the application scenario is thoroughly analyzed and the system
requirements are derived. In general, this conforms to the standard project development process [9], but in our case, we would like to focus on the fault-tolerant
aspects.
Every scenario where a system is deployed has different requirements in terms of
reliability, availability, and lifetime of the system. A satellite system, for example,
has highest demands in all three properties: highest reliability as maintenance is
almost impossible or extremely expensive, highest availability as satellite services
are typically crucial (e.g., GPS system) for the users, and a long lifetime as they
cannot easily be replaced.
It is obvious that these extreme requirements can only be met with a rigorous
design and implementation phase. However, the higher the required level of reliability and availability, the higher is also the cost. In fact, the cost for the design and
implementation of a dependable system rises exponentially with the degree of
required dependence.
It is thus very important not to design a system that is as reliable as possible, as
such a system would be extremely expensive, but to derive from the applications
and the usage scenarios of such a system the required level of reliability.
As shown above, fault tolerance is the main method of increasing the dependability of a system, it is therefore necessary to keep during all phases of software
development (specification, design, implementation, testing, deployment, and
maintenance) fault tolerance in mind.
An example: humans are not perfect, this is a well-known fact. Therefore, a
system does always have some internal faults, for example, software implementation bugs that are also not identifiable with verification tools. Individual applications should therefore be programmed in a defensive way, which allows a software
component, for example, to detect invalid data or invalid software behavior that
does not correspond to its specification.
Even though this might be an obvious way of how to write programs, it is
nevertheless a method of fault tolerance. And it is only of use, if this principle is
used in all applications and of course the runtime system itself.
System testing needs another word of explanation. In the standard development
process, using automatic test cases or manual testing by an engineer tests the
software and hardware. Testing of fault tolerance proves to be a challenge due to
66
6 System Software Support for Hardware Deficiency…
dependency organization.
6.1 System Software Life Cycle Versus Fault Tolerance
The system life cycle of an FT system can be split into two phases: first, the
development phase (design and implementation) and second the operational phase
(deployment, operating use, and maintenance).
The first, and in terms of an FT system almost most important part is the design
phase. In this phase, the application scenario is thoroughly analyzed and the system
requirements are derived. In general, this conforms to the standard project development process [9], but in our case, we would like to focus on the fault-tolerant
aspects.
Every scenario where a system is deployed has different requirements in terms of
reliability, availability, and lifetime of the system. A satellite system, for example,
has highest demands in all three properties: highest reliability as maintenance is
almost impossible or extremely expensive, highest availability as satellite services
are typically crucial (e.g., GPS system) for the users, and a long lifetime as they
cannot easily be replaced.
It is obvious that these extreme requirements can only be met with a rigorous
design and implementation phase. However, the higher the required level of reliability and availability, the higher is also the cost. In fact, the cost for the design and
implementation of a dependable system rises exponentially with the degree of
required dependence.
It is thus very important not to design a system that is as reliable as possible, as
such a system would be extremely expensive, but to derive from the applications
and the usage scenarios of such a system the required level of reliability.
As shown above, fault tolerance is the main method of increasing the dependability of a system, it is therefore necessary to keep during all phases of software
development (specification, design, implementation, testing, deployment, and
maintenance) fault tolerance in mind.
An example: humans are not perfect, this is a well-known fact. Therefore, a
system does always have some internal faults, for example, software implementation bugs that are also not identifiable with verification tools. Individual applications should therefore be programmed in a defensive way, which allows a software
component, for example, to detect invalid data or invalid software behavior that
does not correspond to its specification.
Even though this might be an obvious way of how to write programs, it is
nevertheless a method of fault tolerance. And it is only of use, if this principle is
used in all applications and of course the runtime system itself.
System testing needs another word of explanation. In the standard development
process, using automatic test cases or manual testing by an engineer tests the
software and hardware. Testing of fault tolerance proves to be a challenge due to
66
6 System Software Support for Hardware Deficiency…
