and runtime system are responsible for the recovery point creation, which is part of
this research.
Using [20, 25, 26] as a guideline, we present here a short survey of possible
backward recovery mechanisms only, as this can be applied in a general way to any
program without application or programmer intervention.
A recovery point RP is at minimum the states of the participating processes [20,
26]. In other papers, the term checkpoint is used instead of recovery point. We
prefer, however, to use the term “recovery point” for stored information used the
recovery from faults and “checkpoint” the point in time when the hardware is
checked for faults.
Recovery points are always stored on a stable storage [27], which means that the
storage survives any system fault without corruption. Such a stable storage can
either be memory, a hard disk, or better flash disk or be located on a physically
separated device.
In the process of recovery point creation, the problem of inter-process consistency must be taken into account because the rollback of a failed process can force
non-failed processes to rollback as well for the purpose of restoring program state
consistency.
Let us have an example: In a system, with two processes a and b, a sends
messages to process b. If b fails and has to rollback, it might rollback to a point in
time before the arrival of the last message of a.
To fix this inconsistency, process a has to rollback as well even though it did not
fail. This effect is called rollback propagation. If badly implemented, the rollback
propagation can continue up to the initial software state, which is called domino
effect [28, 29]. The domino effect often occurs if processes create their recovery
points independent of each other.
Figure 8.2 shows a simple example of the domino effect. Two threads P and Q
that communicate with message-passing run concurrently and create their respective recovery points as shown in the figure.
In case of a recovery, the system now tries to find a consistent state space which
does not exist in this example except for the initial state. The reason is simple: The
global state is only consistent if the threads receive only messages that were sent. If
R Q 3 is rolled back, P receives a message from Q that Q never sent, and therefore P
must be rolled back as well. The process continues until the initial state is reached.
A rollback line [28] is defined as the most recent consistent set of process
recovery points. If the recovery points of all processes are consistent, the domino
effect cannot happen, as it is only necessary to recover the system up to the last
Q
P
Fig. 8.2 Example of domino
effect
8.2 Overview of Existing Backward Recovery Techniques
115
Précédent

- 128/315

Suivant