point when they remove the RP message from the message queue. The message
queues must be reliable and FIFO for this to work.
Another approach sends the latest recovery point number together with the
actual message. If the latest RP number of the receiver is smaller than the received
number, the receiver generates a new recovery point [35].
In a distributed system, synchronized clocks can be used to synchronize the RP
creation. The systems generate their RPs at a specific predefined interval and wait
for a predefined interval (clock inaccuracy) before continuation [40]. In a distributed system, the communication properties (reordering, delay, reliability) play
an important role. We ignore these problems, as we don’t want to go into further
details about recovery points in distributed systems.
All introduced approaches force all threads to create a recovery point. In real
systems, however, threads or more specifically programs often exclusively communicate in related groups. The fact can be exploited as Koo and Toueg [41] did.
They introduced a checkpoint initiator that partitions the threads in groups that then
individually create their recovery points.
These groups can be further limited if the type of message is taken into consideration. Only messages sent from one thread to the other introduce dependencies
between the threads, if they contain shared data. Janakiraman et al. [42] therefore
record a dependency only if a message containing shared memory that was modified since the last recovery point is sent from one thread to the other. This approach
progressed by introducing page stamps [43], resulting in further reduced
dependencies.
Communication-induced recovery pointing basically generates a recovery point
whenever inter-process communication occurs. A checkpoint needs to be generated
if modified data is sent to another thread. Incremental recovery pointing is often
used in this approach to limit recovery point size. Obviously, the number of performed recovery points is linear to the number of occurring inter-process communications, i.e., this approach is not well suited for communication-heavy
applications. An example of this approach is [42].
The coordinated recovery point approach has one significant drawback. When
programs heavily communicate with the outside world using input–output schemes
(I/O), this approach is not well suited, as I/O does not fit directly this scheme and
extra application-dependent efforts are required.
8.2.2.1 Limiting Checkpoint Size
The simplest approach to generating a recovery point is to store the whole memory
area the program has access to. However, this can be very expensive if the program
is large or needs a lot of data.
We present here two existing approaches of how to reduce the recovery point
size. Li [44] introduced a page-based approach that uses the memory management
unit of the computer to identify memory pages that were written to (copy on write)
since the last recovery point was taken.
118
8 Recovery Preparation
queues must be reliable and FIFO for this to work.
Another approach sends the latest recovery point number together with the
actual message. If the latest RP number of the receiver is smaller than the received
number, the receiver generates a new recovery point [35].
In a distributed system, synchronized clocks can be used to synchronize the RP
creation. The systems generate their RPs at a specific predefined interval and wait
for a predefined interval (clock inaccuracy) before continuation [40]. In a distributed system, the communication properties (reordering, delay, reliability) play
an important role. We ignore these problems, as we don’t want to go into further
details about recovery points in distributed systems.
All introduced approaches force all threads to create a recovery point. In real
systems, however, threads or more specifically programs often exclusively communicate in related groups. The fact can be exploited as Koo and Toueg [41] did.
They introduced a checkpoint initiator that partitions the threads in groups that then
individually create their recovery points.
These groups can be further limited if the type of message is taken into consideration. Only messages sent from one thread to the other introduce dependencies
between the threads, if they contain shared data. Janakiraman et al. [42] therefore
record a dependency only if a message containing shared memory that was modified since the last recovery point is sent from one thread to the other. This approach
progressed by introducing page stamps [43], resulting in further reduced
dependencies.
Communication-induced recovery pointing basically generates a recovery point
whenever inter-process communication occurs. A checkpoint needs to be generated
if modified data is sent to another thread. Incremental recovery pointing is often
used in this approach to limit recovery point size. Obviously, the number of performed recovery points is linear to the number of occurring inter-process communications, i.e., this approach is not well suited for communication-heavy
applications. An example of this approach is [42].
The coordinated recovery point approach has one significant drawback. When
programs heavily communicate with the outside world using input–output schemes
(I/O), this approach is not well suited, as I/O does not fit directly this scheme and
extra application-dependent efforts are required.
8.2.2.1 Limiting Checkpoint Size
The simplest approach to generating a recovery point is to store the whole memory
area the program has access to. However, this can be very expensive if the program
is large or needs a lot of data.
We present here two existing approaches of how to reduce the recovery point
size. Li [44] introduced a page-based approach that uses the memory management
unit of the computer to identify memory pages that were written to (copy on write)
since the last recovery point was taken.
118
8 Recovery Preparation
