18.3 Resilience and Recoverability in Networked System
Previous section focused on clarification of PRE-smart design for computer systems, the same approach can be applied to networks, and make the latter more
suitable for real-time and safety-critical applications. First of all, is important to
establish the difference between a computer system and Distributed Computer
System (DCS), and this can be summarized in the following [3–5]:
(a) Redundancy in networks is already a fact
(b) Latency of thread impact for network cannot be avoided
(c) The propagation of thread impact for network is flooding-like
Clear that DCS to be PRE-smart and fit new requirements of RT and safety-critical
applications and should deal with (b) and point (c) above. Two concept or principles here might be useful, they are: ASAP and ALAP: stop a threat as soon as
possible (ASAP), and handle it as local as possible (ALAP). Also, as it is described
in full details in [3–5] malfunctions should be handled hierarchically, from hardware up to system level support of software recovery and system, reconfiguration.
In turn, reconfiguration of hardware in DCS is essential when permanent fault
took place. Again, ALAP concept should prevail over other design solutions as
wasting of hardware resources in volume is not feasible. All these arguments and
analysis are published in two books [2–4] and here we just summarize elements that
fit the purpose of making DCS fit for new applications.
Thinking about point (a) above we can see that Recoverability can be achieved in
DCS applying scheme of application redundancy (Fig. 18.1). Let us consider in a
bit more details how it can be implemented in DCS. Figure 18.2 presents a
hypothetic segment of network topology with incoming and internal connections.
Incoming and out-coming edges are represented with arrows.
By looking at Fig. 18.2. one might conclude that structural redundancy of the
topology is considerable and the application of the above topology for real-time and
safety-critical applications is straightforward.
Note here that any thread appeared inside the segment might cause serious
troubles because it might propagate very rapidly and cause irreversible damage to
the elements or the structure of DCS as a whole.
Note that here threads signify faults which could be permanent or just malfunctions of software or hardware components [1–5].
Also, a recoverability for networks requires more efforts and extending of GAFT
introduced in [12], and further developed and explained in [1–7, 8, 9]. These
actions, specifically, are:
(a) Find where threads propagate
(b) Make an estimation of the damages
(c) Stop propagation
(d) Find the source of the threads
18.3 Resilience and Recoverability in Networked System
255
Précédent

- 262/315

Suivant