Chapter 14
Hardware: The ERRIC Architecture
Abstract There is no doubt, system software support of computer resilience is
incomplete when hardware does not have essential properties required to implement
it. Luckily, hardware with required support of recoverability and fault detection is
available and described in full details in [1–3]. Here, we briefly describe available
hardware with attention to implementation support of PRE properties (performance
reliability and energy) requirements.
14.1 Processor Architecture
The main principle used in the design of the ERRIC processor is simplicity. The
instruction set as well as the implementation of the processor is reduced to the
absolute minimum required to support general-purpose computing.
This allows the careful controllable introduction of redundancy features that
allow the processor to detect malfunctions during the execution of the instruction
itself, abort the current instruction, and re-execute it transparently to software.
In addition, no pipelining or caches are used which greatly simplifies the processor design which in turn simplifies the implementation of any fault-tolerant
features.
Antola [4] proved that the overheads necessary to make a CPU fully fault
tolerance might easily exceed 100% which in fact corresponds to at least duplication. Thus, to this huge overhead, we strive to keep the redundancy level needed
to achieve fault tolerance as low as possible.
In comparison to other architectures such as CISC processors, ERRIC differs in
the following features: one addressing mode, hardwired design (no microcode),
regular instruction set, and few and simple instructions.
The difference of ERRIC to RISC processors is not as big as to CISC processors,
but, for example, the number of instructions is still smaller and the instruction
complexity lower (no multiply, etc.).
Figure 14.1 shows a simplified version of the instruction execution of the
ERRIC processor. The execution steps correspond to the ones in other RISC
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_14
197
Précédent

- 208/315

Suivant