Introduction for Second Edition
When in 1989 an anonymous reviewer commented on my short paper that “this
classification should be extended to description of distributed systems,” (Yet
another approach to classification of redundancy, CIM IMEKO Symposium 1990,
Helsinki, pp. 117–124) I was really excited, because people in the research community were thinking much deeper and wider than myself (- I had just defended my
Ph.D.).
Further, fault tolerance was migrating to dependability (Jean Claude Laprie was
an indisputable authority and expert in this domain, see more www.springer.com/
gb/book/9783709191729, which later emerged as the concept of resilience.
In principle, all these new properties had concrete reasoning and meaning behind
them: when something erroneous happens, any system of our design should be able
to cope with the problem. Options vary, as well as circumstances and area of
application, thus:
– If it stops the error propagating and freezes in a safe state, it is fail-stop, or
fail-safe;
– If it can cope with permanent faults inside the system, it is a fault-tolerant
system;
– When it continues with reduced functionality, it is graceful degradation;
– If it is designed with attention having been paid to reliability, availability, and
maintenance or serviceability, it is dependable system;
– If it is capable of tolerating obstacles caused by internal and external factors and
can spring back, recover, and continue, then a system can be considered as
resilient.
There are two major ways to achieve any of the properties mentioned above: at
system level or at local level (technological). Obviously, any reasonable combination of both levels is also welcome. We do not want to repeat our papers and
books (https://www.springer.com/gb/book/9783319150680, https://www.springer.
com/gb/book/9783319468129) but to incorporate into the second edition any significant progress that has emerged.
vii
Précédent

- 6/315

Suivant