Triple modular redundancy [7] is a very simple solution and perhaps the most
implemented approach in RT safety-critical applications. However, the area and
therefore also the power consumption penalty is large (higher than 200%).
Since the malfunction sensitivity of sequential and combinational logic is
increasing dramatically [1, 2] and the multiple bit upset ratio of SRAM is also
rising, the previous mechanisms by themselves will not provide the necessary
reliability for safety-critical systems.
2.2 Single Event Effects and Other Deviations
Semiconductor devices experience single event upsets in two major forms: in the
form of destructive effects which result in permanent degradation or even complete
failure of the device and therefore also affecting functionality (permanent fault), and
in the form of nondestructive effects, causing no permanent damage (malfunctions).
However, in this sense when something irregular affects the system the question
must be answered which of the two cases it is. How to deal with this deviation is the
part of the HW/SSW design.
In terms of events, the most common are single event upsets and multiple cell
upsets, which both belong to the single event category.
Single-Bit Upset is a single event upset or single-bit upset, meaning, one event
produces a single-bit error. This type of error is very common in SRAM.
Multiple Cell Upset is a multiple bit upset for one event regardless of the
location of the multiple bits. For example, an FPGA where one routing bit gets an
impact from a high energetic particle, affecting several memory positions. Hence,
multiple cell upsets involve both upsets, the ones that can be corrected by error
correction codes and those which cannot with reasonable overhead.
Multiple Bit Upset is a subset of multiple cell upsets. It is a multiple bit upset for
one event that affects several bits in the same word. This type of deviation cannot be
corrected by error correction codes with reasonable overhead. However, it is possible to partially avoid multiple bit upsets by using specific layout design of
memory cells.
Growing density of logic elements of wafer technology, miniaturization of
manufacturing processes, and high clock speeds will inevitably increase rate of
so-called intermittent faults.
Inevitable variations in the manufacturing process will induce these faults at
higher rate; moreover, impact of these faults will be lasting up to several seconds
[2], increasing complexity of recovery. In contrast to external faults such as radiation, intermittent faults are triggered by internal events such as voltage, temperature, and timing variations.
Figure 2.2 illustrates such an intermittent fault.
2.1 Introduction
9
Précédent

- 25/315

Suivant