The computation process consists of the sequential execution of program segments where a recovery point is made after the execution of every segment. The
actual successful segment execution time is considered as a random process.
Even though a program segment in case of no failure is executed in time t with T
as the mean segment execution time, the randomness is required as occurring
malfunctions or permanent faults and resulting recovery operations need additional
time, depending on the executed recovery steps.
We assume that the checking intervals and the intervals between two recovery
points are the same, as it is possible to combine hardware testing and recovery point
creation into one step [2, 4]. Despite this change, the computation model is the same
as the one used in the MLR approach explained above.
We now want to compare the three recovery algorithms (linear, dichotomous,
and our modified linear) by their mean execution time to execute program segment
k. We also want to take into account the possible latencies of the hardware
malfunctions.
In the analysis of the MLR algorithm, we included malfunctions and permanent
faults. In this chapter, we assume that only malfunctions can be recovered successfully, and these also only if the influence of any occurring malfunction has
stopped before any recovery actions are initiated.
Permanent faults always lead to unsuccessful recovery. We also assume that no
additional errors occur during recovery. The following events might happen during
the execution of a program segment:
– H c : No malfunction occurred during the execution of the segment and it has
been successfully executed in the first run.
– H m : A malfunction occurred during the execution of a program segment, the
segment has been correctly identified, and is successfully finished after recovery
step m. The number of total recovery steps is M.
– H pf : A malfunction occurred during the execution of a program segment, the
segment has been correctly identified but all K recovery steps did not result in
successful recovery. Recovery is achieved after reconfiguration and program
restart.
– H m (L): A malfunction occurred during the execution of a segment but the latent
malfunction caused the failed segment to be incorrectly identified.
– H pf (L): A malfunction occurred during the execution of a program segment, but
the latent fault caused the failed element to be incorrectly identified and all K
iterations did not result in successful recovery. Recovery is only achieved after
reconfiguration and program restart.
The fault distribution during the execution of a segment is assumed to have a
Poisson distribution with the parameter k = k 1 + k 2 , with k 1 as the rate of malfunctions and k 2 as the rate of permanent faults [6, 7].
154
10 Recovery Algorithms: An Analysis
Précédent

- 167/315

Suivant