As faults in hardware can either be detected by the hardware itself (mismatch
and voters) or software-based checks, both the hardware and software need access
to the syndrome registers.
Hardware checks. If the hardware detects a fault, the hardware updates the syndrome and raises an interrupt. The software can then handle the interrupt and further
diagnose the hardware. If multiple faults are detected, the principle of growing core
must be applied, mode details are presented in [10].
The same is valid if consecutive faults occur during fault handling. The syndrome interrupt should therefore be priority-based and reentrant. For instance, a
fault in a peripheral device must have a lower priority than a fault in the ALU.
When the software finished handling the fault, it must clear the handled syndrome
bit.
Software checks. Software-based hardware checks use specific hardware properties
for checking, i.e., clearly defined hardware behavior or software redundancy when
checking for faults in memory. Although these checks are initiated by software,
they can also trigger a hardware-checking scheme which then signals the fault via
interrupt and syndrome. An example for this is memory scrubbing that triggers a
latent fault (modified memory cell) on triplicated memory.
As the interrupt controller itself can also be affected by a fault, which could
result in a nonworking syndrome, interrupt, the software must periodically poll the
fault syndrome to check for faults.
The ERA instruction set (ISA) (see Table 14.1), which the simulator uses, does
not include special purpose instructions to access the syndrome. Using some of the
special registers instead of general purpose ones for the syndrome would contradict
the design principles of the instruction set.
Thus, the preferred way to operate and access the syndrome that avoids changing
the ISA is the use of the I/O memory lines to a reserved fixed location in the address
space.
In hardware, the syndrome scheme could be implemented in a triplicated
scheme, either connected upstream to the memory controller or in parallel.
Regardless of the triplicated implementation, only one syndrome register set is
visible to the software. An error in one of the syndrome registers is therefore
corrected by the hardware without software intervention.
However, it seems useful that the runtime is aware of errors in the syndrome
itself for monitoring purposes. In addition, the monitoring of all occurring errors
indicated by the syndrome provides useful information for a potential contingency
plan (e.g., setting the fault tolerance of the system to a higher level in case of recent
particle impacts).
Although both hardware- and software-based checks are required for full possible fault coverage, the hardware implemented schemes of detection fault and
signal to the rest of the system are clearly the superior approach.
The hardware never clears a bit in the fault syndrome on its own, as it cannot
know when the software finished handling the fault; therefore, the software must
clear all handled faults manually before it jumps back from the syndrome interrupt.
7.3 System Monitoring of Checking Process: A Syndrome
93
and voters) or software-based checks, both the hardware and software need access
to the syndrome registers.
Hardware checks. If the hardware detects a fault, the hardware updates the syndrome and raises an interrupt. The software can then handle the interrupt and further
diagnose the hardware. If multiple faults are detected, the principle of growing core
must be applied, mode details are presented in [10].
The same is valid if consecutive faults occur during fault handling. The syndrome interrupt should therefore be priority-based and reentrant. For instance, a
fault in a peripheral device must have a lower priority than a fault in the ALU.
When the software finished handling the fault, it must clear the handled syndrome
bit.
Software checks. Software-based hardware checks use specific hardware properties
for checking, i.e., clearly defined hardware behavior or software redundancy when
checking for faults in memory. Although these checks are initiated by software,
they can also trigger a hardware-checking scheme which then signals the fault via
interrupt and syndrome. An example for this is memory scrubbing that triggers a
latent fault (modified memory cell) on triplicated memory.
As the interrupt controller itself can also be affected by a fault, which could
result in a nonworking syndrome, interrupt, the software must periodically poll the
fault syndrome to check for faults.
The ERA instruction set (ISA) (see Table 14.1), which the simulator uses, does
not include special purpose instructions to access the syndrome. Using some of the
special registers instead of general purpose ones for the syndrome would contradict
the design principles of the instruction set.
Thus, the preferred way to operate and access the syndrome that avoids changing
the ISA is the use of the I/O memory lines to a reserved fixed location in the address
space.
In hardware, the syndrome scheme could be implemented in a triplicated
scheme, either connected upstream to the memory controller or in parallel.
Regardless of the triplicated implementation, only one syndrome register set is
visible to the software. An error in one of the syndrome registers is therefore
corrected by the hardware without software intervention.
However, it seems useful that the runtime is aware of errors in the syndrome
itself for monitoring purposes. In addition, the monitoring of all occurring errors
indicated by the syndrome provides useful information for a potential contingency
plan (e.g., setting the fault tolerance of the system to a higher level in case of recent
particle impacts).
Although both hardware- and software-based checks are required for full possible fault coverage, the hardware implemented schemes of detection fault and
signal to the rest of the system are clearly the superior approach.
The hardware never clears a bit in the fault syndrome on its own, as it cannot
know when the software finished handling the fault; therefore, the software must
clear all handled faults manually before it jumps back from the syndrome interrupt.
7.3 System Monitoring of Checking Process: A Syndrome
93
