12.7 Concurrency: Further Steps: Fault-Tolerant
Interactors
This section is dealing with the description of principle of fault tolerance for
semaphores that extend a condition proposed by Dijkstra during NATO school
presentations about losing process out of the critical section.
Real problem of safety-critical systems is exactly opposite—dealing with process behavior when this process is in its critical section.
An approach is obvious—integer semaphore (not Boolean) must be applied and
recovery points such as in 1987 organized for each process entering into the critical
section.
Also—if you read carefully Dijkstra assumptions, his synchronization assumes
that noncritical section behavior of a process is not important. And this is a pretty
naïve assumption, but reconfigurability of hardware was reasonably described in [7,
8] and in runtime system description section there is no room for further analysis of
this statement.
But! and this but is big: the behavior of a process in critical section CANNOT be
ignored. The assumption that eventually process will leave its critical section is not
strong enough: synchronization is about to maximize performance, to kick out a
process out of CS by a timer—10 million years later—is not a solution—see the
statement about time is not an option.
The question if system resilience and functions of runtime systems monitors
come unanswered:
What if a process that own critical section at the moment “died” or “hang up”
in there for an unlimited time? All other processes will wait and die eventually.
What can we do?
That is how an idea of faul-tolerant semaphore came to life.
Thus, the fault-tolerant semaphore means that any change of process condition
due to internal fault must automatically lead to
• release a critical section and
• release all system resources that were used.
Further, it implements fault-tolerant semaphore—if we are using the algorithm of
Fig. 12.7 we will require a change of variable “turn” by reducing the number of
processes that will be permitted to apply for access to the critical section after the
fault of a process has happened. Thus, for N scheduled processes monitor of
interruptions should reduce the number of “applicants” to N–1.
To avoid fast degradation of the system due to reconfigurability of hardware in
this case one has to (eventually, when the workload is not critical) to check the
health of a process that was self-declared dead. This is because, for example of
natural radiation impact after which malfunctions in the real-time systems now lasts
up to for 200–350 ms. Book [7], papers [2, 3, 8] and some chapters of this book
might explain details of reconfiguration handling.
190
12 Proposed Runtime System Structure …
Précédent

- 202/315

Suivant