Chapter 16
ERRIC Reliability
Abstract The key property of our design is resilience. So far we mostly covered
system software methods and schemes to achieve or support it. Previous Chaps. 14
and 15 explained briefly what hardware (processor) should possess (including
reduction on functions and limited in architecture options) to be able to implement
resilience in the most efficient way. Here, we intend to analyze what we have
achieved in hardware design in terms of malfunction tolerance—attempting to use
heavy artillery of system software as less as possible, making dirty work of fault
detection and determination (malfunction or permanent for hardware).
16.1 ERRIC Reliability Analysis
Having introduced the reliability analysis of malfunction tolerance in Sect. 4.6.1,
we apply it here to the current ERRIC processor prototype. Current estimations
show a fault coverage of 100% for single-bit upsets (faults that affect one bit in the
current operation).
We do not consider multiple faults although it is expected that with growths of
die density, the amount of multiple faults will be higher than the amount of single
faults. This is subject to future work.
The actual implementation of the ERRIC processor shows that hardware overheads in the range of 12% for checking and 3% for recovery [1, 2].
Using these two values, we can estimate the Mean Time To Failure (MTTF) of
this system and also the reliability as a function of time.
The MTTF nf and reliability over time (P nf (t)) of ERRIC without fault tolerance
are given by Eq. 4.4 and Eq. 4.3, respectively. For the MTTF ft and P ft (t) of ERRIC,
we can apply Eq. 4.8 and Eq. 4.7, respectively. We do not use Eq. 4.11 as this
equation is not directly applicable to a concrete system.
We recapitulate the equation pairs here reflecting them with figures:
P mf ¼ e
Àð1 þ kÞk pf 1 t
ð16:1Þ
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_16
215
Précédent

- 224/315

Suivant