Chapter 9
Recovery: Searching and Monitoring
of Correct Software States
Abstract The last of the three GAFT processes is called recovery and recovery
monitoring. After the detection of an error and possible reconfiguration, the last step
is recovering the software, which means that the effect of the error on the software
must be eliminated. In line with the previous chapters and [1–6], the recovery
consists of restoring the last recovery point and continuing the processing. But is
this really sufficient? What if latent faults exist in the system and manifest themselves in the system but trigger some detection schemes an arbitrary time later?
Assuming this reasonable and unpleasant sequence of events, it becomes clear that
just restoring data and program from the last stored recovery point is not enough.
We have to admit that we do not have any guarantee that fault is now eliminated:
even when hardware is restored or even reconfigured—we have erroneous states of
software recorded in recovery points. Thus, we have to consider the recovery
process itself and analyze which classic algorithms are applicable and fit the purpose of efficient recovery. We introduce and analyze three recovery algorithms that
are able to ensure successful recovery by iteratively go through all stored recovery
points.
9.1 Recovery as a Process
If a fault is detected in the system, appropriate recovery actions must be executed to
eliminate the fault and to recover the software.
The software recovery process consists of the following actions:
• interrupt processing,
• identify the found fault,
• eliminate the found fault,
• check the integrity of the system, and
• recover software.
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_9
141
Précédent

- 154/315

Suivant