The charged particle passes through the semiconductor material leaving an
ionized track behind leaving sufficient energy in the circuit to have an effect on a
localized area of the electronic device.
Single event upsets can occur either through the impact of primary particles
(e.g., direct ionization from galactic cosmic rays or solar particles) or by the secondary particles generated after the strike (indirect ionization). It affects many
different types of devices, designed using various technologies resulting in data
corruption, high current conditions, and transient disturbances. If such single event
upsets are not handled well, unwanted functional interrupts and catastrophic failures
could take place.
However, single event upset is not fully representative anymore. The number of
multiple bits affected by a single event was relatively small using previous silicon
technologies. Modern technologies are much more vulnerable and single event can
affect multiple bits, so-called Multiple Bits Upsets (MBU). The single event rate
affecting multiple bits is expected to increase in the coming years [2].
Traditionally, two different types of techniques have been used to mitigate
upsets: fault avoidance and fault tolerance techniques. Fault avoidance techniques,
usually at the device level such as silicon on insulator or hardened memory cells,
usually involve IC process changes. These techniques have drawbacks in terms of
cost, chip area, and speed of operation.
Several low-level fault-tolerant techniques for memories and processor registers
have been used to mitigate upsets: error detection codes and error correction codes.
Hamming codes [4] have been largely employed to detect and correct single-bit
upsets. However, these techniques are vulnerable to the increasing multiple bit
upset ratio.
Reed–Solomon codes [5] can correct a large number of multiple faults but are not
suitable for hardware implementation in terms of design complexity, additional
memory required (area), and inherent latency (performance).
At a higher level, circuit-level mitigation techniques with some amount of redundancy, such as N-modular redundancy and voting [6], are frequently used.
Fig. 2.1 Comparison of the
SRAM bit soft error rate
(malfunctions) [8]
8
2 Hardware Faults
ionized track behind leaving sufficient energy in the circuit to have an effect on a
localized area of the electronic device.
Single event upsets can occur either through the impact of primary particles
(e.g., direct ionization from galactic cosmic rays or solar particles) or by the secondary particles generated after the strike (indirect ionization). It affects many
different types of devices, designed using various technologies resulting in data
corruption, high current conditions, and transient disturbances. If such single event
upsets are not handled well, unwanted functional interrupts and catastrophic failures
could take place.
However, single event upset is not fully representative anymore. The number of
multiple bits affected by a single event was relatively small using previous silicon
technologies. Modern technologies are much more vulnerable and single event can
affect multiple bits, so-called Multiple Bits Upsets (MBU). The single event rate
affecting multiple bits is expected to increase in the coming years [2].
Traditionally, two different types of techniques have been used to mitigate
upsets: fault avoidance and fault tolerance techniques. Fault avoidance techniques,
usually at the device level such as silicon on insulator or hardened memory cells,
usually involve IC process changes. These techniques have drawbacks in terms of
cost, chip area, and speed of operation.
Several low-level fault-tolerant techniques for memories and processor registers
have been used to mitigate upsets: error detection codes and error correction codes.
Hamming codes [4] have been largely employed to detect and correct single-bit
upsets. However, these techniques are vulnerable to the increasing multiple bit
upset ratio.
Reed–Solomon codes [5] can correct a large number of multiple faults but are not
suitable for hardware implementation in terms of design complexity, additional
memory required (area), and inherent latency (performance).
At a higher level, circuit-level mitigation techniques with some amount of redundancy, such as N-modular redundancy and voting [6], are frequently used.
Fig. 2.1 Comparison of the
SRAM bit soft error rate
(malfunctions) [8]
8
2 Hardware Faults
