Chapter 18
Distributed Systems: Maximizing
Resilience
Igor Schagaev
Abstract We claim that system redundancy (natural or artificial, deliberately
introduced) should be applied for the purposes of Performance, Reliability and
Energy efficiency or “PRE-smartness”. A process of an algorithm of application of
available redundancy for PRE-smartness is proposed showing that it is further
extension generalized algorithm of fault tolerance, when property of fault detection
is replaced as property of PRE-smartness. We present a system level implementation steps of PRE-smartness. Using ITACS LTD forward and backward tracing
algorithms applied for distributed computer system we demonstrate that efficiency
(reliability as a whole, especially availability; performance of detection and recovery) grow up to the level when real-time applications of distributed system
grows substantially. Ans estimation of reliability gain is modeled and demonstrated
suggesting a policy of implementation of PRE-smartness as a permanent on-going
process required for the system successful functioning.
18.1 Introduction
Desperation management is one of the concepts we use in everyday life for effective
handling of our conditions using available resources. Obviously, organization,
hierarchy and structure of these resources within any system, not only in our body
or brain (knowledge IS the resource!) can either help or complicate their utilization.
For networked computers or distribute computer systems further (DCS) as a
whole we can consider as resources the following (we are not pretending to make a
federal level of classification here):
– topology of the system
– elements—nodes (routers, switches etc.)
– rules—algorithms of behavior control for of the system as a whole
– routing algorithms
– network protocols
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_18
249
Distributed Systems: Maximizing
Resilience
Igor Schagaev
Abstract We claim that system redundancy (natural or artificial, deliberately
introduced) should be applied for the purposes of Performance, Reliability and
Energy efficiency or “PRE-smartness”. A process of an algorithm of application of
available redundancy for PRE-smartness is proposed showing that it is further
extension generalized algorithm of fault tolerance, when property of fault detection
is replaced as property of PRE-smartness. We present a system level implementation steps of PRE-smartness. Using ITACS LTD forward and backward tracing
algorithms applied for distributed computer system we demonstrate that efficiency
(reliability as a whole, especially availability; performance of detection and recovery) grow up to the level when real-time applications of distributed system
grows substantially. Ans estimation of reliability gain is modeled and demonstrated
suggesting a policy of implementation of PRE-smartness as a permanent on-going
process required for the system successful functioning.
18.1 Introduction
Desperation management is one of the concepts we use in everyday life for effective
handling of our conditions using available resources. Obviously, organization,
hierarchy and structure of these resources within any system, not only in our body
or brain (knowledge IS the resource!) can either help or complicate their utilization.
For networked computers or distribute computer systems further (DCS) as a
whole we can consider as resources the following (we are not pretending to make a
federal level of classification here):
– topology of the system
– elements—nodes (routers, switches etc.)
– rules—algorithms of behavior control for of the system as a whole
– routing algorithms
– network protocols
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_18
249
