Chapter 2
Hardware Faults
Abstract This chapter explains hardware faults, their origins, dependency on
technology used, and some known solutions of fault toleration using error correction codes and redundancy hardware schemes. Shown that with growing density of
hardware there is a risk of multiple temporary fault which grows at order of
magnitude prime concern for designers of new computer systems for safety-critical
application. Hardware faults occur due to natural phenomena such as ionized
radiation, variations in the manufacturing process, vibrations, etc. We present in this
chapter a short introduction to hardware faults, show the typical fault types and
patterns, and also give examples of how to deal with these faults.
2.1 Introduction
Transient faults are a huge concern in silicon-based electronic components such as
SRAM, DRAM, microprocessors, and FPGA. Those are devices that have a
well-documented history of transient faults mainly caused by energetic particles.
An important obstacle for safety-critical systems to achieve maximum reliability
is represented by the susceptibility of those systems to faults produced by radiation.
In addition, as manufacturing technologies evolve, the effects of ionizing radiation
become a primary concern.
Due to the reduction in size of the transistors and the reduction in critical charge
of logic circuits, the susceptibility of technologies to information corruption is
increasing [1–3]. A concrete example for this is given in Fig. 2.1 [1], indicating that
the malfunction rate (soft error rate) is massively increasing in processor logic for
decreasing manufacturing sizes.
Collisions of particles with sensitive regions of the semiconductor can change
stored information in ROM/RAM and lead to logic errors, for example, in processor
circuits. In this context, Single Event Upsets (SEU) are effects induced by the strike
of a single energetic particle (ion, proton, electron, neutron, etc.) in the
semiconductor.
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_2
7
Hardware Faults
Abstract This chapter explains hardware faults, their origins, dependency on
technology used, and some known solutions of fault toleration using error correction codes and redundancy hardware schemes. Shown that with growing density of
hardware there is a risk of multiple temporary fault which grows at order of
magnitude prime concern for designers of new computer systems for safety-critical
application. Hardware faults occur due to natural phenomena such as ionized
radiation, variations in the manufacturing process, vibrations, etc. We present in this
chapter a short introduction to hardware faults, show the typical fault types and
patterns, and also give examples of how to deal with these faults.
2.1 Introduction
Transient faults are a huge concern in silicon-based electronic components such as
SRAM, DRAM, microprocessors, and FPGA. Those are devices that have a
well-documented history of transient faults mainly caused by energetic particles.
An important obstacle for safety-critical systems to achieve maximum reliability
is represented by the susceptibility of those systems to faults produced by radiation.
In addition, as manufacturing technologies evolve, the effects of ionizing radiation
become a primary concern.
Due to the reduction in size of the transistors and the reduction in critical charge
of logic circuits, the susceptibility of technologies to information corruption is
increasing [1–3]. A concrete example for this is given in Fig. 2.1 [1], indicating that
the malfunction rate (soft error rate) is massively increasing in processor logic for
decreasing manufacturing sizes.
Collisions of particles with sensitive regions of the semiconductor can change
stored information in ROM/RAM and lead to logic errors, for example, in processor
circuits. In this context, Single Event Upsets (SEU) are effects induced by the strike
of a single energetic particle (ion, proton, electron, neutron, etc.) in the
semiconductor.
© Springer Nature Switzerland AG 2020
I. Schagaev et al., Software Design for Resilient Computer Systems,
https://doi.org/10.1007/978-3-030-21244-5_2
7
