that might be connected with computer state changes including hardware faults,
time-outs, or other interruptions, including interaction with other processes.
Then, if of the loop “hangs” due to problems within hardware an arrival of
another signal able to break loop execution might become visible—hardware state
change reflects immediately within the control construction and we are not using
“brute force” of conventional waiting or wasting resources of another type.
FT as a new feature of embedded systems, implemented on the level of system
software should be analyzed in some detail. First, let us consider fault tolerance as a
sequence of steps as introduced by GAFT in chapter.
When we defined FT as GAFT it became possible to investigate how system
software should be involved to realize this algorithm. Then, all required features
and mechanisms of system software to support fault tolerance of embedded systems
can be derived from GAFT.
There are several processes and functions in the system software to provide fault
tolerance, Fig. 6.6.
Research in the area of program recovery after a hardware fault has occurred is
known for 40 years. Google search shows millions of related links, including some
even from Microsoft. But none have been found that is concerned with the actual
semantics of checking and recovery at the system software at programming language level and runtime system level.
Regretfully, checking and recovery sometimes are used as synonyms. Checking
is the process of analyzing the hardware state to prove an existence (or absence) of
faults from predefined classes. In turn, recovery is the process of preparation and
storing of several states of the hardware aimed at being able to recover from any
predefined and arranged state or states.
Both processes, including recovery schemes, are important and unavoidable
elements of GAFT. Their combination improves the actual reliability of RT FT
systems and the efficiency of solutions taken at least at the model level.
Various schemes for program recovery after a hardware fault were introduced
and described in [6]. Several methods of correct recovery searching with supporting
hardware are described in [7, 8] including scheme of recovery point formation;
algorithms of recovery and their efficiency, and management and efficient supportive schemes.
It is worth to mention that recovery points might be arranged at the level of tasks
(that are mutually dependent either with data or control or both, or interrelated by
conversations). These recovery algorithms become much more complex and almost
prohibitively slow for RT applications. This kind of system recovery might be a
Fig. 6.6 Processes and functions for system software to implement FT
6 System Software Support for Hardware Deficiency…
65
time-outs, or other interruptions, including interaction with other processes.
Then, if of the loop “hangs” due to problems within hardware an arrival of
another signal able to break loop execution might become visible—hardware state
change reflects immediately within the control construction and we are not using
“brute force” of conventional waiting or wasting resources of another type.
FT as a new feature of embedded systems, implemented on the level of system
software should be analyzed in some detail. First, let us consider fault tolerance as a
sequence of steps as introduced by GAFT in chapter.
When we defined FT as GAFT it became possible to investigate how system
software should be involved to realize this algorithm. Then, all required features
and mechanisms of system software to support fault tolerance of embedded systems
can be derived from GAFT.
There are several processes and functions in the system software to provide fault
tolerance, Fig. 6.6.
Research in the area of program recovery after a hardware fault has occurred is
known for 40 years. Google search shows millions of related links, including some
even from Microsoft. But none have been found that is concerned with the actual
semantics of checking and recovery at the system software at programming language level and runtime system level.
Regretfully, checking and recovery sometimes are used as synonyms. Checking
is the process of analyzing the hardware state to prove an existence (or absence) of
faults from predefined classes. In turn, recovery is the process of preparation and
storing of several states of the hardware aimed at being able to recover from any
predefined and arranged state or states.
Both processes, including recovery schemes, are important and unavoidable
elements of GAFT. Their combination improves the actual reliability of RT FT
systems and the efficiency of solutions taken at least at the model level.
Various schemes for program recovery after a hardware fault were introduced
and described in [6]. Several methods of correct recovery searching with supporting
hardware are described in [7, 8] including scheme of recovery point formation;
algorithms of recovery and their efficiency, and management and efficient supportive schemes.
It is worth to mention that recovery points might be arranged at the level of tasks
(that are mutually dependent either with data or control or both, or interrelated by
conversations). These recovery algorithms become much more complex and almost
prohibitively slow for RT applications. This kind of system recovery might be a
Fig. 6.6 Processes and functions for system software to implement FT
6 System Software Support for Hardware Deficiency…
65
