However, if the fault detection mechanism identified the wrong unit as the faulty
one, then the MLR algorithm results in a loop as the still present faulty unit in the
system triggers the fault again after recovery.
Algorithms that eliminate permanent hardware faults usually consist of three
successively executed phases (denoted by A, B, and C, respectively): A: determination of the fault type, B: elimination of permanent hardware faults (hardware
reconfiguration), and C: recovery of the correct program state and continuation of
the computing process. These three phases are of course a subset of GAFT.
From GAFT, we know that the permanent hardware fault elimination depends
on the results of the fault type detection. During fault type detection, the program
segment or instruction is rerun, during which the checking logic or test program
identified the error.
For the recovery of the correct program state, we recommend to use the MLR
algorithm which requires less time to establish the last correct program state than
other known algorithms of the same family. The time required to execute a full
recovery is the sum of the time required to perform the three steps fault type
detection, permanent hardware fault elimination, and recovery of the correct program state. As no computation is performed while the system recovers, the minimization of the time required for full recovery is very important to reach maximum
availability and to keep the system response time as short as possible.
Let us now examine the recovery process when the MLR algorithm is used for
the fault type determination. Assume that a given system has a hardware failure
which has been identified by the checking logic or checking procedures in the
program segment j.
As long as the permanent fault is not eliminated, the computing process even if
recovered by the current correct RP will remain corrupted. At the same time, if the
fault affected the system operation during the first generation of an RP, then, CS and
CS’ will match in reruns of the next program segment due to the fault being
permanent. If a program fragment preceding the occurrence of a fault is rerun, a
non-eliminated fault will corrupt the result of the computation and CS and CS’ will
not match. The same checksum mismatch will occur when the recovery depth is
further incremented.
Let us now consider a program execution with n program segments. Depending
on the time of the failure manifestation, the match–mismatch sequences have the
form, shown in Figs. 9.2, 9.3 and 9.4.
By examining the sequence patterns, it is possible to decide whether the fault is a
malfunction or a permanent fault. The first (rightmost) mm in the sequences of
Figs. 9.2 and 9.3, which may be preceded by one or more m’s, corresponds to an
RP generated before the failure has been manifested.
This RP contains, therefore, the last correct state of the process. For the cases
shown in Figs. 9.2 and 9.3, the process can be recovered and continued from this
RP after the hardware failure is fixed. If no mismatch is observed during all n
recovery steps (Fig. 9.4), the failure occurred before the leftmost RP was created.
If earlier recovery points exist, the recovery process simply continues, if after all
n recovery steps the first RP is reached, then the only way to recover is to reboot the
9.2 Modified Linear Algorithm
147
one, then the MLR algorithm results in a loop as the still present faulty unit in the
system triggers the fault again after recovery.
Algorithms that eliminate permanent hardware faults usually consist of three
successively executed phases (denoted by A, B, and C, respectively): A: determination of the fault type, B: elimination of permanent hardware faults (hardware
reconfiguration), and C: recovery of the correct program state and continuation of
the computing process. These three phases are of course a subset of GAFT.
From GAFT, we know that the permanent hardware fault elimination depends
on the results of the fault type detection. During fault type detection, the program
segment or instruction is rerun, during which the checking logic or test program
identified the error.
For the recovery of the correct program state, we recommend to use the MLR
algorithm which requires less time to establish the last correct program state than
other known algorithms of the same family. The time required to execute a full
recovery is the sum of the time required to perform the three steps fault type
detection, permanent hardware fault elimination, and recovery of the correct program state. As no computation is performed while the system recovers, the minimization of the time required for full recovery is very important to reach maximum
availability and to keep the system response time as short as possible.
Let us now examine the recovery process when the MLR algorithm is used for
the fault type determination. Assume that a given system has a hardware failure
which has been identified by the checking logic or checking procedures in the
program segment j.
As long as the permanent fault is not eliminated, the computing process even if
recovered by the current correct RP will remain corrupted. At the same time, if the
fault affected the system operation during the first generation of an RP, then, CS and
CS’ will match in reruns of the next program segment due to the fault being
permanent. If a program fragment preceding the occurrence of a fault is rerun, a
non-eliminated fault will corrupt the result of the computation and CS and CS’ will
not match. The same checksum mismatch will occur when the recovery depth is
further incremented.
Let us now consider a program execution with n program segments. Depending
on the time of the failure manifestation, the match–mismatch sequences have the
form, shown in Figs. 9.2, 9.3 and 9.4.
By examining the sequence patterns, it is possible to decide whether the fault is a
malfunction or a permanent fault. The first (rightmost) mm in the sequences of
Figs. 9.2 and 9.3, which may be preceded by one or more m’s, corresponds to an
RP generated before the failure has been manifested.
This RP contains, therefore, the last correct state of the process. For the cases
shown in Figs. 9.2 and 9.3, the process can be recovered and continued from this
RP after the hardware failure is fixed. If no mismatch is observed during all n
recovery steps (Fig. 9.4), the failure occurred before the leftmost RP was created.
If earlier recovery points exist, the recovery process simply continues, if after all
n recovery steps the first RP is reached, then the only way to recover is to reboot the
9.2 Modified Linear Algorithm
147
