8.2.1 Uncoordinated Recovery Points
Uncoordinated recovery points approach belongs to the first implemented recovery
techniques [28]. This technique sets no constraints on the creation of recovery
points with the advantage that the thread can initiate a recovery point creation when
the state space is small which results in small recovery points.
The implementation is also quite simple. The downside, however, is the susceptibility to the domino effect, which reduces the practicability of this approach
considerable. It might also happen that a thread creates an RP that is not part of the
global system state and wastes therefore processing time.
As the global consistent system state is in general not known, every thread must
maintain several RPs, which need to be garbage collected over time to regain space.
During recovery, it is necessary to determine a global consistent system state to
which the system is then recovered. When a thread sends messages to another, it
adds the current checkpoint number to the message. The receiver then stores this
number locally. In case of a rollback, these numbers are used to calculate a
dependency graph showing the latest global consistent state [38].
8.2.2 Coordinated Recovery Points
In contrast to the uncoordinated recovery points approach, threads are synchronized
if coordinated recovery points are used. The main idea is that if threads are synchronized between each other, the state is always system-wide consistent.
Therefore, every thread needs to keep only one stored recovery point. One of the
earliest approaches [39] implements blocking recovery point coordination approach
that uses two-phase commit to force all threads to create a recovery point at the
same time.
Synchronization then is achieved by blocking all threads at the same time
(message broadcast). This accompanied by flushing of all communication channels
and creating the recovery points. The disadvantage of this approach is obvious:
While the recovery points are taken, the whole system is blocked and unable to
react to external events, i.e., become nonresponsive.
Between the recovery point creations, there is no additional overhead involved in
this approach and the total overhead is adjustable by the recovery point interval.
The non-blocking recovery point coordination approaches use other synchronization mechanisms that do not force the system to stop. An early approach [29]
solves the problem of inconsistency by creating a recovery point prior to sending a
message to another thread.
Extension to that use markers that are sent from one thread that generated a
recovery point to all others. These generate their checkpoint and broadcast the
marker further. This is an example of a non-blocking recovery point approach, as
the threads do not immediately stop program execution but create the recovery
8.2 Overview of Existing Backward Recovery Techniques
117
Précédent

- 130/315

Suivant