If, however, the testing procedures still detect the error, the fault in the system is
a long-lasting fault. Thus, the system is recovered to RP i
’ and the processing is
resumed. When the program reaches RP i , i.e., the point in the program execution
when RP i was created, the recovery point mechanism creates a new recovery point
RP i ’ and the checksum CS i ’. The checksums CS i and CS i ’ are now compared. In
case of CS i = CS i ’, the fault is still present, in case of CS i 6 ¼ CS i ’ the program
calculations changed, i.e., the fault is eliminated.
This procedure is executed recursively until the system is recovered or restarted
if recovery completely failed. During recovery, no new recovery points RPs are
created, of course.
Let us assume that a malfunction occurred at time t a and that the duration (m) of
its immediate effect on the computing process exceeds the execution time of one or
more program segments between the recovery points:
Dm [ xt f
ð9:1Þ
Here x (x > 1) is the number of program fragments, and t f is the execution time of
each fragment. In case of timely even distributed recovery points, t f is constant.
During the interval Δm (Fig. 9.1), both the task and the generation of the RPs are
executed by faulty hardware.
During interval R, starting at the point in time the effect of the malfunction ends,
to the time the malfunction is detected by some check or test facilities, the program
is executed on correctly working hardware but with incorrect data that was modified
by the malfunction. Obviously, the recovery points that are generated during this
time contain also incorrect data.
Is it possible now to find the last correct RP? There is an answer actually. At
every iteration of the recovery procedure, the recovery depth is increased by one
step or segment, the program executed and the resulting checksums CS and CS’
compared. Obviously, CS and CS’ match in the interval R—it is indicated by the
letters m (match) in Fig. 9.1 on the right. Of special interest is the interval Δm.
During the generation of the recovery points, the content of some of the RPs
might have been influenced directly by the effect of the malfunction on the program
execution and by the use of incorrect data affected by the malfunction.
During recovery, the program, which was affected by the malfunction (period
Δm), is now run on properly working hardware (either tested or checked by
hardware schemes). Therefore, the program is repeated under different conditions
with incorrect data on correctly operating hardware. Consequently, the pairs CS and
CS’ will not match in the interval Δm, as indicated in Fig. 9.1 by the letters mm
(mismatch).
When the system is recovered to the last recovery point before the malfunction
occurred and the program segment executed, the resulting checksums CS and CS’
match, since the hardware and software that was used during the real data processing and during the recovery were in the same state. The task has thus been
recovered and can be continued [3, 7, 8, 11].
9.2 Modified Linear Algorithm
145
Précédent

- 158/315

Suivant