In the optimal case, an RP includes the CPU registers and the affected memory
areas such as the program code and data sections, the stack, the heap, and the
relevant operating system data structures. If operating system structures and CPU
registers are not included in the RP, the RP and the respective CS cover exactly the
application, which might lead to unsuccessful recovery as the fault affected kernel
data structures (i.e., I/O data).
Special care must be taken when choosing the checksum generator. A simple
modulo 2 adder, for example, can detect any single fault in the program or data and
will result in different CS.
However, for multiple faults, this is no longer guaranteed. CRCs or Hamming
codes provide a better fault coverage and easy hardware implementation. Details
about CRC analysis can be found in [12], for Hamming codes in [13, 14].
Cryptographic hash values are the alternative, which are specifically designed for
collision avoidance and do not rely on specific error patterns. The downside,
however, is that it’s not possible to make any guarantees about fault coverage, and
their implementation in hardware is complex and expensive.
If, as proposed in [15], analysed in [16] and developed in [7, 8], the RPs cover
only current program data, a malfunction corrupting the program code or execution
might not result in different RP, CS pairs. Such CS generation impairs the diagnostic power properties of checksum comparison, i.e., it may not be possible to
recognize the type of the fault and to determine the correct RP.
If a fault already exists at boot up time, the RP and CS and the comparison of the
CSs of the first and following runs allow to identify such a fault as a permanent fault
since the fault will never disappear. When fault exists ate the boot up time the RP
and CS identify such fault as permanent.
9.2.2 MLR Execution in Case of Permanent Faults
The detection of malfunctions and permanent faults uses the same techniques. In
both cases, the checking circuit or checking process identifies the fault and the
effect of the fault is seen in corrupted data or program code and also misbehaving
program execution. This results eventually in the generation of one or more RPs
that include corrupt data and program state. However, recovery depends on the fault
type as explained Sect. 4.4.
Thus, the MLR algorithm can in the case of permanent faults not directly be
applied. The system must eliminate the permanent fault first or go into a degraded
state, by excluding the faulty component. The MLR algorithm can therefore only
correctly restore the system state if the fault type is known.
If we assume that in case of a permanent fault the faulty unit is replaced with a
fault-free one, resulting in the elimination of the permanent fault, then the MLR
algorithm successfully recovers the system.
146
9 Recovery: Searching and Monitoring …
areas such as the program code and data sections, the stack, the heap, and the
relevant operating system data structures. If operating system structures and CPU
registers are not included in the RP, the RP and the respective CS cover exactly the
application, which might lead to unsuccessful recovery as the fault affected kernel
data structures (i.e., I/O data).
Special care must be taken when choosing the checksum generator. A simple
modulo 2 adder, for example, can detect any single fault in the program or data and
will result in different CS.
However, for multiple faults, this is no longer guaranteed. CRCs or Hamming
codes provide a better fault coverage and easy hardware implementation. Details
about CRC analysis can be found in [12], for Hamming codes in [13, 14].
Cryptographic hash values are the alternative, which are specifically designed for
collision avoidance and do not rely on specific error patterns. The downside,
however, is that it’s not possible to make any guarantees about fault coverage, and
their implementation in hardware is complex and expensive.
If, as proposed in [15], analysed in [16] and developed in [7, 8], the RPs cover
only current program data, a malfunction corrupting the program code or execution
might not result in different RP, CS pairs. Such CS generation impairs the diagnostic power properties of checksum comparison, i.e., it may not be possible to
recognize the type of the fault and to determine the correct RP.
If a fault already exists at boot up time, the RP and CS and the comparison of the
CSs of the first and following runs allow to identify such a fault as a permanent fault
since the fault will never disappear. When fault exists ate the boot up time the RP
and CS identify such fault as permanent.
9.2.2 MLR Execution in Case of Permanent Faults
The detection of malfunctions and permanent faults uses the same techniques. In
both cases, the checking circuit or checking process identifies the fault and the
effect of the fault is seen in corrupted data or program code and also misbehaving
program execution. This results eventually in the generation of one or more RPs
that include corrupt data and program state. However, recovery depends on the fault
type as explained Sect. 4.4.
Thus, the MLR algorithm can in the case of permanent faults not directly be
applied. The system must eliminate the permanent fault first or go into a degraded
state, by excluding the faulty component. The MLR algorithm can therefore only
correctly restore the system state if the fault type is known.
If we assume that in case of a permanent fault the faulty unit is replaced with a
fault-free one, resulting in the elimination of the permanent fault, then the MLR
algorithm successfully recovers the system.
146
9 Recovery: Searching and Monitoring …
