Chapter 8
Recovery Preparation
Abstract In the last section, we showed how hardware integrity of a computing
system can be efficiently ensured using hardware-checking schemes and system
software testing procedures and their sequences. However, to recover from faults, it
is necessary to eliminate the effects the error had on the computation, i.e., the
software code and data space. In GAFT, this corresponds to preparation for recovery. We now want to show how software has to be organized to be able own
recovery or in other words, we want to revise different strategies how software can,
after the detection of an error, ensure that the error did not affect the software state,
or if this cannot be ensured, what precautions software has to conduct to be able to
re-establish a correct software state. First, we revise the state of the art and then
introduce a new technology and show its power and limitations. In the next step, we
will show how hardware can assist software in the process of recovery preparation.
For all generic approaches to recovery preparation, so-called stable storage, a
nonvolatile, reliable, and fast storage is needed. If no direct hardware support is
available, stable storage must be implemented in software. We will present a
possible software implementation of such a stable storage.
8.1 Runtime System Support for Fault Tolerance
and Reconfigurability
The runtime system, often also called kernel of an operating system, provides
low-level software functionality such as threading, inter-process communication,
and memory management. These three tasks are the minimum functionality every
runtime system must provide; otherwise, an operating system cannot be implemented [1]. Since we analyze design and develop of a single (not distributed),
embedded computer system with properties of fault tolerance, reconfigurability, etc.
and abilities of implementation of these properties at the level of hardware and
especially system software we do not pursue the analysis if distributed systems.
Some of aspects of implementation of fault tolerance in distributed systems are
covered in our previous work [2].
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_8
111
Précédent

- 124/315

Suivant