First, a diagnosis procedure signals a fault via an interrupt or a system call. The
system must immediately stop processing data and handle the fault to prevent the
fault from spreading in the system.
If the fault identification mechanism that found the fault is not powerful enough
to diagnose the fault type, it is necessary to run further diagnostic routines, starting
from repeating of instruction [4, 5]. In specific, every hardware driver must provide
a diagnosis procedure that identifies the fault type. If the checking is fully
hardware-based, this software procedure might only read the fault registers
(Sect. 7.3), if it is software-based, it might run some more sophisticated tests. Based
on the found result, the checking procedure eliminates the fault. The assumption
here is that every device driver provides a catalogue of fault elimination mechanisms, for example, “power off” and “power on” the faulty device and rerun the test
to show the absence of faults.
If the device is still faulty, it should be excluded from the current working set
and turned off (Sect. 7.4). The hardware itself should now be fault-free again;
however, software might have been affected by the fault. Dependent on which level
the checking is performed (instruction, procedure, task, etc.) the software is
recovered differently.
Recovery on the instruction level As the fault was detected during the instruction
execution, no software recovery is necessary. However, in the case of permanent
faults, a hardware reconfiguration is still required to eliminate the fault.
Recovery on the procedure level and above In this case, it is assumed that faults
are detected in the time frame of procedure execution. It is therefore sufficient to
perform a rollback based on the last stored recovery point. In fact, a rollback is the
inverse function of creating a recovery point, with the difference that the data is not
copied from the memory and stored on the recovery point storage, but the data is
read from this storage and copied back to the memory.
Prior to that, recalculating the checksum and comparing it to the stored one must
ensure the consistency of the recovery point. If the comparison succeeds, the recovery point is considered as valid; otherwise, it is discarded, and the previous
recovery point is used for recovery.
In fact, recovery on the level of procedures does not differ from recovery on a
higher level. First, the content of the recovery point is restored, then the stack and
base pointer, and as the last step, the processing is resumed by jumping to the
restored instruction pointer.
From an implementation perspective, the code, which is executed to restore the
recovery point, should be executed from ROM, using a reserved memory area for
temporary data so that recovered data cannot overwrite data of the recovery process
itself. In case of Oberon using module variables only (no heap data structures)
would be sufficient.
However, the recovery process still needs its own private stack. Simulating a
private heap in a statically allocated buffer would be of course also possible, but not
elegant and unnecessary complicated.
142
9 Recovery: Searching and Monitoring …
system must immediately stop processing data and handle the fault to prevent the
fault from spreading in the system.
If the fault identification mechanism that found the fault is not powerful enough
to diagnose the fault type, it is necessary to run further diagnostic routines, starting
from repeating of instruction [4, 5]. In specific, every hardware driver must provide
a diagnosis procedure that identifies the fault type. If the checking is fully
hardware-based, this software procedure might only read the fault registers
(Sect. 7.3), if it is software-based, it might run some more sophisticated tests. Based
on the found result, the checking procedure eliminates the fault. The assumption
here is that every device driver provides a catalogue of fault elimination mechanisms, for example, “power off” and “power on” the faulty device and rerun the test
to show the absence of faults.
If the device is still faulty, it should be excluded from the current working set
and turned off (Sect. 7.4). The hardware itself should now be fault-free again;
however, software might have been affected by the fault. Dependent on which level
the checking is performed (instruction, procedure, task, etc.) the software is
recovered differently.
Recovery on the instruction level As the fault was detected during the instruction
execution, no software recovery is necessary. However, in the case of permanent
faults, a hardware reconfiguration is still required to eliminate the fault.
Recovery on the procedure level and above In this case, it is assumed that faults
are detected in the time frame of procedure execution. It is therefore sufficient to
perform a rollback based on the last stored recovery point. In fact, a rollback is the
inverse function of creating a recovery point, with the difference that the data is not
copied from the memory and stored on the recovery point storage, but the data is
read from this storage and copied back to the memory.
Prior to that, recalculating the checksum and comparing it to the stored one must
ensure the consistency of the recovery point. If the comparison succeeds, the recovery point is considered as valid; otherwise, it is discarded, and the previous
recovery point is used for recovery.
In fact, recovery on the level of procedures does not differ from recovery on a
higher level. First, the content of the recovery point is restored, then the stack and
base pointer, and as the last step, the processing is resumed by jumping to the
restored instruction pointer.
From an implementation perspective, the code, which is executed to restore the
recovery point, should be executed from ROM, using a reserved memory area for
temporary data so that recovered data cannot overwrite data of the recovery process
itself. In case of Oberon using module variables only (no heap data structures)
would be sufficient.
However, the recovery process still needs its own private stack. Simulating a
private heap in a statically allocated buffer would be of course also possible, but not
elegant and unnecessary complicated.
142
9 Recovery: Searching and Monitoring …
