– The software needs direct access to the hardware checksum register to read the
current checksum value.
– The checksum and recovery point device must provide two operating modes:
First, the standard mode, which creates the recovery points and its checksum
and seconds a recovery mode, which generates only the checksum. The operating mode is configurable by the software.
– The checksum generator should have its own dedicated connection to the stable
storage.
– For performance reasons, the stable storage should be directly attached to the
computer so that the processor has direct access for maximum speed during RP
generation and recovery.
– The RP and its associated CS should be generated concurrently to speed up the
RP generation.
9.3 Summary
We presented a short conceptual description of the general recovery procedure
required to achieve resilience, fault tolerance, and reliability.
We then introduced the modified linear recovery algorithm that as we showed is
able to decide whether the recovery was successful or whether the error is still
present in the system by comparing the content of recovery points or their
checksum.
We also showed that the modified linear recovery algorithm could even be used
to distinguish the fault type, i.e., permanent fault or malfunction. Thus, it covers in
fact part of the testing and checking process as well.
Multiple faults do not interfere the modified linear recovery algorithm during
recovery, and the MLR algorithm is able to detect them as separate faults.
The MLR algorithm has one strict requirement, which is repeatability of program on hardware.
References
1. Sogomonian E, Schagaev I (1988) Hardware and software fault tolerance of computer
systems. Avtom I Telemekhanika, 3–39
2. Schagaev I (1989) Computing process recovery algorithms. Avtomat Telemekh (4)
3. Schagaev I (1990) Using software recovery methods for determining the type of hardware
faults. Autom Remote Control 51(3)
4. Schagaev I (2008) Reliability of malfunction tolerance. In: International multi-conference on
computer science and information technology, 2008. IMCSIT 2008, October 2008, pp 733–
737
5. Schagaev I et al (2010) ERA: evolving reconfigurable architecture. In: 11th ACIS
International Conference, June 2010, pp 215–220
9.2 Modified Linear Algorithm
151
current checksum value.
– The checksum and recovery point device must provide two operating modes:
First, the standard mode, which creates the recovery points and its checksum
and seconds a recovery mode, which generates only the checksum. The operating mode is configurable by the software.
– The checksum generator should have its own dedicated connection to the stable
storage.
– For performance reasons, the stable storage should be directly attached to the
computer so that the processor has direct access for maximum speed during RP
generation and recovery.
– The RP and its associated CS should be generated concurrently to speed up the
RP generation.
9.3 Summary
We presented a short conceptual description of the general recovery procedure
required to achieve resilience, fault tolerance, and reliability.
We then introduced the modified linear recovery algorithm that as we showed is
able to decide whether the recovery was successful or whether the error is still
present in the system by comparing the content of recovery points or their
checksum.
We also showed that the modified linear recovery algorithm could even be used
to distinguish the fault type, i.e., permanent fault or malfunction. Thus, it covers in
fact part of the testing and checking process as well.
Multiple faults do not interfere the modified linear recovery algorithm during
recovery, and the MLR algorithm is able to detect them as separate faults.
The MLR algorithm has one strict requirement, which is repeatability of program on hardware.
References
1. Sogomonian E, Schagaev I (1988) Hardware and software fault tolerance of computer
systems. Avtom I Telemekhanika, 3–39
2. Schagaev I (1989) Computing process recovery algorithms. Avtomat Telemekh (4)
3. Schagaev I (1990) Using software recovery methods for determining the type of hardware
faults. Autom Remote Control 51(3)
4. Schagaev I (2008) Reliability of malfunction tolerance. In: International multi-conference on
computer science and information technology, 2008. IMCSIT 2008, October 2008, pp 733–
737
5. Schagaev I et al (2010) ERA: evolving reconfigurable architecture. In: 11th ACIS
International Conference, June 2010, pp 215–220
9.2 Modified Linear Algorithm
151
