rollback line. Under the assumption that the system is a fail-stop system and faults
manifest themselves immediately, it’s just necessary to keep that last recovery line.
Every earlier recovery point can be freed.
The classic system model which is used when analyzing recovery points is a
fail-stop system [29], i.e., they either work according to specification or crash (stop
working) without corrupting data, especially memory.
Most of the work which has already been done in this field is based on
message-passing systems, which can either be distributed system with a network
interconnection, but also single multiprocessor systems where messages are passed
via shared memory or other local interconnection facility. In fact, the single
multi-process message-passing system is just a special case of a distributed message
passing, and thus the same algorithms can be applied.
The commonly used approaches for creating recovery points are either implemented on the level of the system or on the level of tasks. The former is also called
global checkpointing [30, 31]. The global checkpointing assumes that the recovery
points are created when all other operations of the system stop, i.e., the whole
system operations are “frozen” for creating a recovery point of the whole system.
This approach has, of course, the huge disadvantage that the system must be
completely stopped while the snapshot is made. In this sense, recoverability action
reduces reliability—long-range preparation of recovery from future malfunctions or
permanent faults reduce point availability by adding to maintenance time and
reducing operational time.
Recovery points on the level of tasks have already been extensively covered in
literature and can be classified as follows:
Uncoordinated recovery points Every task creates a recovery point whenever it
suits the task [28];
Coordinated recovery points Chandy and Lamport [29] introduced coordinated
recovery points by saving a system-wide consistent state. Some processes [30] save
their state in a coordinated way;
Message logging In these systems, recovery points of threads are created independently from each other. Logging all messages between threads retains consistency between threads. In case of a failure, the threads are rolled back and the
communication is replayed and thus consistency restored.
The systems assume piecewise determinism [31], which requires that all determinants [32], i.e., all non-deterministic event, can be identified and logged.
Typically, messages between threads and I/O data are determinants. Examples for
message logging systems are [33–37].
These mechanisms are all application transparent, as we have mentioned the
application-dependent techniques before. Recovery on the level of procedures is
covered in Sect. 8.3.
116
8 Recovery Preparation
manifest themselves immediately, it’s just necessary to keep that last recovery line.
Every earlier recovery point can be freed.
The classic system model which is used when analyzing recovery points is a
fail-stop system [29], i.e., they either work according to specification or crash (stop
working) without corrupting data, especially memory.
Most of the work which has already been done in this field is based on
message-passing systems, which can either be distributed system with a network
interconnection, but also single multiprocessor systems where messages are passed
via shared memory or other local interconnection facility. In fact, the single
multi-process message-passing system is just a special case of a distributed message
passing, and thus the same algorithms can be applied.
The commonly used approaches for creating recovery points are either implemented on the level of the system or on the level of tasks. The former is also called
global checkpointing [30, 31]. The global checkpointing assumes that the recovery
points are created when all other operations of the system stop, i.e., the whole
system operations are “frozen” for creating a recovery point of the whole system.
This approach has, of course, the huge disadvantage that the system must be
completely stopped while the snapshot is made. In this sense, recoverability action
reduces reliability—long-range preparation of recovery from future malfunctions or
permanent faults reduce point availability by adding to maintenance time and
reducing operational time.
Recovery points on the level of tasks have already been extensively covered in
literature and can be classified as follows:
Uncoordinated recovery points Every task creates a recovery point whenever it
suits the task [28];
Coordinated recovery points Chandy and Lamport [29] introduced coordinated
recovery points by saving a system-wide consistent state. Some processes [30] save
their state in a coordinated way;
Message logging In these systems, recovery points of threads are created independently from each other. Logging all messages between threads retains consistency between threads. In case of a failure, the threads are rolled back and the
communication is replayed and thus consistency restored.
The systems assume piecewise determinism [31], which requires that all determinants [32], i.e., all non-deterministic event, can be identified and logged.
Typically, messages between threads and I/O data are determinants. Examples for
message logging systems are [33–37].
These mechanisms are all application transparent, as we have mentioned the
application-dependent techniques before. Recovery on the level of procedures is
covered in Sect. 8.3.
116
8 Recovery Preparation
