13. Kaegi-Trachsel T, Gutknecht J (2008) Minos—the design and implementation of an
embedded real-time operating system with a perspective of fault tolerance. International
Multiconference on IMCSIT 2008, 20–22 October 2008, pp 649–656
14. Fabry RS (1974) Capability-based addressing. Commun ACM 17:403–412
15. Schagaev I (1990) Using software recovery facilities for determining the type of hardware
faults. Autom and Remote Control 51(3)
16. McCluskey E et al (2002) Control-flow checking by software signatures. IEEE Trans Reliab
51(1):111–122
17. Schagaev I (1989) Computing process recovery algorithms. Avtomat Telemekh 4
18. Oh N, Mitra S, McCluskey (2002) Error detection by diverse data and duplicated instructions.
IEEE Trans Comput 51(2):180–199
19. McCluskey E et al (2002) Error detection by duplicated instructions in superscalarprocessors.
IEEE Trans Reliab 51(1):63–75
20. Sogomonyan E, Schagaev I (1988) Hardware and software for fault-tolerant computing
systems. Autom Remote Control 49:129–151
21. McCluskey E et al (2000) Dependable computing and online testing in adaptive and
configurable systems. IEEE Des Test Comput 17(1):29–41
22. Mukherjee S et al (2002) Detailed design and evaluation of redundant multi-threading
alternatives. In: 29th annual international symposium on computer architecture, pp 99–110
23. Dal Cin M et al (1993) Fault tolerance in distributed shared memory multiprocessors. In:
Parallel computer architectures: theory, hardware, software, applications. Springer, London,
pp 31–48
24. Candea G, Kawamoto S, Fujiki Y, Greg Friedman G, Fox A (2004) Microreboot: a technique
for cheap recovery. In: Proceedings of the 6th conference on symposium on operating systems
design & implementation, vol 6. USENIX Association, Berkeley, CA, USA, p 6
25. Deconinck G et al (1993) Survey of backward error recovery techniques for multicomputers
based on checkpointing and rollback. Int J Model Simul 18:262–265
26. Elnozahy E et al (2002) A survey of rollback-recovery protocols in message-passing systems
27. Lampson BW (1981) Atomic transactions. In: Distributed systems—architecture and
implementation, an advanced course. Springer, London, pp 246–265
28. Randell B (1975) System structure for software fault tolerance. IEEE Trans Softw Eng 1:
220–232
29. Lamport L et al (1985) Distributed snapshots: determining global states of distributed
systems. ACM Trans Comput Syst 3:63–75
30. Attig N, Sander V (1993) Automatic checkpointing of NQS batch jobs on CRAY unicos
systems. In: Proceedings of the cray user group meeting, pp 250–255
31. Strom R, Yemini S (1985) Optimistic recovery in distributed systems. ACM Trans Comput
Syst 3:204–226
32. Lorenzo A, Keith M (1996) Trade-offs in implementing causal message logging protocols. In:
15th ACM symposium on principles of distributed computing, PODC ’96. ACM, New York,
NY, USA, pp 58–67
33. Borg A, Baumbach J, Glazer S (1983) A message system supporting fault tolerance. In:
Proceedings of the ninth ACM symposium on operating systems principles, SOSP ’83. ACM,
New York, NY, USA, pp 90–99
34. Strom R, Bacon D, Yemini S (1988) Volatile logging in n-fault-tolerant distributed systems.
In: Digest of papers eighteenth international symposium on fault-tolerant computing,
FTCS-18, pp 44–49
35. Elnozahy E, Zwaenepoel W (1992) Manetho: transparent roll back-recovery with low
overhead, limited rollback, and fast output commit. IEEE Trans Comput 41(5):526–531
36. Johnson D, Zwaenepoel W (1987) Sender-based message logging. In: Digest of papers: 17
annual international symposium on fault-tolerant computing. IEEE Computer Society,
pp 14–19
138
8 Recovery Preparation
Précédent

- 151/315

Suivant