5. In case of a malfunction, the event is logged and the program execution
resumed. Logging the events is important as an accumulation of malfunctions in
a module or a specific memory location could hint to a permanent failure in the
near future.
6. In case of a permanent fault, the current memory configuration is extracted from
the syndrome, and the next degradation state is calculated according to the
application needs and predefined degradation tables.
7. The new calculated memory configuration is written to the syndrome registers
and the power of the faulty unit disabled. In some configuration transitions, the
software has to adapt to the new situation and recover after excluding the faulty
unit but before including the new one.
8. Software clears the handled fault in the syndrome and continues processing by
returning from the interrupt.
Some of the presented transitions in Sect. 7.3.4 need software intervention to fully
recover from a permanent fault and to establish a new working software state. We
show here a list of all situations where software support is needed.
Adding/Replacing a module of an already populated bank If a memory module
of a duplicated or triplicated memory mode is replaced by another module, the
memory content must be replicated to the new module before it is included in the
configuration. If this is not done, every memory access would yield an error.
A small loop in software, using only registers is sufficient to perform the copy
without modifying the memory during the copy operation. To have access to the
spare module, the software configures the syndrome to attach the spare module
temporarily to a free memory bank. After the copy operation, the spare module can
be included in the working set. These actions must be performed, for example,
when the system switches from Phase 1 to Phase 2 in Fig. 7.17.
Failed module of bank 1 Bank 1 contains by convention all runtime system data
structures and is therefore critical for system operation. If the memory module of
bank 1 fails, the software crashes ultimately.
The only way to treat this is by hard resetting the system either via a hardware
watchdog or a software initiated power cycle. The built-in self-test of the hardware
automatically identifies the failed module and configures another still working
module to bank 1. The runtime system can then restart all critical applications.
Failed module of another bank Assuming that the runtime system structures do
not cover more than one module, the runtime system can simply free all resources
used by all applications and restart them appropriately.
Instead of graceful degradation, the software can also decide to “regenerate”, i.e.,
going from a mode with less redundancy to a mode with higher redundancy. The
inclusion of a spare module corresponds to Point 1 in the list above; if an already
used module is moved to another bank, software has to release all data structures
residing on that module in case it is in nonredundant use and repopulate it with data
according to Point 1.
102
7 Testing, Checking, and Hardware Syndrome
resumed. Logging the events is important as an accumulation of malfunctions in
a module or a specific memory location could hint to a permanent failure in the
near future.
6. In case of a permanent fault, the current memory configuration is extracted from
the syndrome, and the next degradation state is calculated according to the
application needs and predefined degradation tables.
7. The new calculated memory configuration is written to the syndrome registers
and the power of the faulty unit disabled. In some configuration transitions, the
software has to adapt to the new situation and recover after excluding the faulty
unit but before including the new one.
8. Software clears the handled fault in the syndrome and continues processing by
returning from the interrupt.
Some of the presented transitions in Sect. 7.3.4 need software intervention to fully
recover from a permanent fault and to establish a new working software state. We
show here a list of all situations where software support is needed.
Adding/Replacing a module of an already populated bank If a memory module
of a duplicated or triplicated memory mode is replaced by another module, the
memory content must be replicated to the new module before it is included in the
configuration. If this is not done, every memory access would yield an error.
A small loop in software, using only registers is sufficient to perform the copy
without modifying the memory during the copy operation. To have access to the
spare module, the software configures the syndrome to attach the spare module
temporarily to a free memory bank. After the copy operation, the spare module can
be included in the working set. These actions must be performed, for example,
when the system switches from Phase 1 to Phase 2 in Fig. 7.17.
Failed module of bank 1 Bank 1 contains by convention all runtime system data
structures and is therefore critical for system operation. If the memory module of
bank 1 fails, the software crashes ultimately.
The only way to treat this is by hard resetting the system either via a hardware
watchdog or a software initiated power cycle. The built-in self-test of the hardware
automatically identifies the failed module and configures another still working
module to bank 1. The runtime system can then restart all critical applications.
Failed module of another bank Assuming that the runtime system structures do
not cover more than one module, the runtime system can simply free all resources
used by all applications and restart them appropriately.
Instead of graceful degradation, the software can also decide to “regenerate”, i.e.,
going from a mode with less redundancy to a mode with higher redundancy. The
inclusion of a spare module corresponds to Point 1 in the list above; if an already
used module is moved to another bank, software has to release all data structures
residing on that module in case it is in nonredundant use and repopulate it with data
according to Point 1.
102
7 Testing, Checking, and Hardware Syndrome
