Thus, the applications are responsible for reobtaining the capabilities they need.
A capability fault handler, which is raised in case of a failed capability can restore
the state of a capability, e.g., based on a stored checkpoint.
However, recovery is application specific and is therefore semi-transparent. In
case of used resources such as frame buffers, an additional cleanup handler is
required.
Minix 3: Minix 3 [5] is another microkernel operation system that claims to be fault
tolerant. However, the fault tolerance in this operating system is limited to fault
containment by isolated processes and component restart in case of failure.
A reincarnation server restarts the failed component and notifies applications about
the restart. The state of the component is lost after the restart, thus the application
programmer interaction is required to resume processing.
CuriOS: CuriOS [6] is a capability-based client–server operating system that stores
client-related state space, so-called Server State Regions, on the server in distinct
memory areas which are only accessible by the server if it serves a request by a
client. In case of an error on the server (C++ exception), the client-related state
space is still available an thus the server can transparently resume operation. This
model is designed for misbehaving drivers or other software related errors, but
cannot help in case of hardware malfunctions or permanent error that corrupt the
state space.
EROS: EROS [7] is also a capability-based Operating System that employs
checked recovery points of the whole system to recover from faults. Fault detection
is based on failures of tasks, and recovery is done by reloading the last global
snapshot.
The assumption, in this case, is that the last recovery point is consistent.
Recovery from permanent faults is not supported, nor reconfiguration of the
hardware. The recovery point consistency checks, as we see it, just ensure the
consistency of the recovery point but not of the system at the point of the recovery
point creation.
Table 13.1 lists the main steps of GAFT and states which of the steps are
implemented by the above introduced operating systems. This table clearly shows
that from a GAFT point of view no OS except for ERA implements all steps.
Especially, the steps Fault-type determination and Hardware reconfiguration are
not natively supported by any OS except ERA. For recovery, most of the OS just
restart a failed component and component failure is usually detected by assertions,
exceptions or watchdogs.
Obviously, masked faults or malfunctions that do not trigger one of the before
mentioned schemes are not detected. Only L4ReAnimator supports a checkpointing
scheme, which can recover the system to an earlier state. Overall, the ERA concept
seems superior in terms of fault tolerance to the other approaches.
The other operating systems, especially the microkernel-based ones have better
fault containment than ERA and can also deal with malicious programs.
194
13 Proposed Runtime System Versus Existing Approaches
A capability fault handler, which is raised in case of a failed capability can restore
the state of a capability, e.g., based on a stored checkpoint.
However, recovery is application specific and is therefore semi-transparent. In
case of used resources such as frame buffers, an additional cleanup handler is
required.
Minix 3: Minix 3 [5] is another microkernel operation system that claims to be fault
tolerant. However, the fault tolerance in this operating system is limited to fault
containment by isolated processes and component restart in case of failure.
A reincarnation server restarts the failed component and notifies applications about
the restart. The state of the component is lost after the restart, thus the application
programmer interaction is required to resume processing.
CuriOS: CuriOS [6] is a capability-based client–server operating system that stores
client-related state space, so-called Server State Regions, on the server in distinct
memory areas which are only accessible by the server if it serves a request by a
client. In case of an error on the server (C++ exception), the client-related state
space is still available an thus the server can transparently resume operation. This
model is designed for misbehaving drivers or other software related errors, but
cannot help in case of hardware malfunctions or permanent error that corrupt the
state space.
EROS: EROS [7] is also a capability-based Operating System that employs
checked recovery points of the whole system to recover from faults. Fault detection
is based on failures of tasks, and recovery is done by reloading the last global
snapshot.
The assumption, in this case, is that the last recovery point is consistent.
Recovery from permanent faults is not supported, nor reconfiguration of the
hardware. The recovery point consistency checks, as we see it, just ensure the
consistency of the recovery point but not of the system at the point of the recovery
point creation.
Table 13.1 lists the main steps of GAFT and states which of the steps are
implemented by the above introduced operating systems. This table clearly shows
that from a GAFT point of view no OS except for ERA implements all steps.
Especially, the steps Fault-type determination and Hardware reconfiguration are
not natively supported by any OS except ERA. For recovery, most of the OS just
restart a failed component and component failure is usually detected by assertions,
exceptions or watchdogs.
Obviously, masked faults or malfunctions that do not trigger one of the before
mentioned schemes are not detected. Only L4ReAnimator supports a checkpointing
scheme, which can recover the system to an earlier state. Overall, the ERA concept
seems superior in terms of fault tolerance to the other approaches.
The other operating systems, especially the microkernel-based ones have better
fault containment than ERA and can also deal with malicious programs.
194
13 Proposed Runtime System Versus Existing Approaches
