For hardware-based checking, we introduced the syndrome a mean of configuring the hardware and signaling detected faults to the software. We discussed
possible hardware reconfiguration scenarios and analyzed in which cases software
has to adapt to the new topology.
Hardware-based checking scheme the so-called syndrome, a hardware controller
that can configure the hardware (power, memory topology) and signal detected
faults to the software. We showed which steps are necessary for the software to
execute in order to identify the fault type (malfunction and permanent error) and in
case of permanent error how software can reconfigure the hardware into a degraded
but still working state.
Based on concrete degradation scenarios, we showed that the software can
exclude faulty memory modules but has to adapt itself in two cases:
• First, if a memory module in a redundant mode is replaced by another one,
software has to repopulate the new module with the data of the working one.
• The second if the main module containing the runtime system data in linear
mode is replaced by another one, the whole system must be rebooted for
recovery.
For a possible hardware implementation of the reconfiguration mechanism, we
propose to use the T-logic embedded in a modified memory controller which also
includes the syndrome logic.
The benefit of such a system is twofold:
• As long as enough hardware elements, e.g., memory modules, are working, the
system can continue processing.
• Hardware can adapt to the requirements of the application: Highly reliable or
more space and more performance as the application has more space available
for processing.
References
1. Kirby W et al (1985) The NMFECC cray time-sharing system. Softw Pract Exper 15(1):87–
103
2. Serlin O (1984) Fault-tolerant systems in commercial applications. Computer C 7(8):19–30
3. Blazewicz J et al (2007) Handbook on scheduling, from theory to applications. Springer,
Berlin, Heidelberg
4. Ingo M (2002) Linux kernel archive. World Wide Web electronic publication, January 03,
2002
5. Bogdanov J, Schagaev I (1990) Sliding slotting diagnosis in multiprocessors. In: IMECO
congress proceedings, pp 141–150
6. Garey M, Johnson D (1979) Computers and in-tractability: a guide to the theory of
NP-completeness. W.H. Freeman and Company
7. Knuth D (1998) The art of computer programming 3. Sorting and searching, vol III.
Addison-Wesley Longman, Amsterdam
7.6 Summary
109
possible hardware reconfiguration scenarios and analyzed in which cases software
has to adapt to the new topology.
Hardware-based checking scheme the so-called syndrome, a hardware controller
that can configure the hardware (power, memory topology) and signal detected
faults to the software. We showed which steps are necessary for the software to
execute in order to identify the fault type (malfunction and permanent error) and in
case of permanent error how software can reconfigure the hardware into a degraded
but still working state.
Based on concrete degradation scenarios, we showed that the software can
exclude faulty memory modules but has to adapt itself in two cases:
• First, if a memory module in a redundant mode is replaced by another one,
software has to repopulate the new module with the data of the working one.
• The second if the main module containing the runtime system data in linear
mode is replaced by another one, the whole system must be rebooted for
recovery.
For a possible hardware implementation of the reconfiguration mechanism, we
propose to use the T-logic embedded in a modified memory controller which also
includes the syndrome logic.
The benefit of such a system is twofold:
• As long as enough hardware elements, e.g., memory modules, are working, the
system can continue processing.
• Hardware can adapt to the requirements of the application: Highly reliable or
more space and more performance as the application has more space available
for processing.
References
1. Kirby W et al (1985) The NMFECC cray time-sharing system. Softw Pract Exper 15(1):87–
103
2. Serlin O (1984) Fault-tolerant systems in commercial applications. Computer C 7(8):19–30
3. Blazewicz J et al (2007) Handbook on scheduling, from theory to applications. Springer,
Berlin, Heidelberg
4. Ingo M (2002) Linux kernel archive. World Wide Web electronic publication, January 03,
2002
5. Bogdanov J, Schagaev I (1990) Sliding slotting diagnosis in multiprocessors. In: IMECO
congress proceedings, pp 141–150
6. Garey M, Johnson D (1979) Computers and in-tractability: a guide to the theory of
NP-completeness. W.H. Freeman and Company
7. Knuth D (1998) The art of computer programming 3. Sorting and searching, vol III.
Addison-Wesley Longman, Amsterdam
7.6 Summary
109
