220
Network-on-Chip
retransmitted. If such a flit arrives correctly, the receiver can deliver to the
network layer, in sequence, all the flits it has buffered. This retransmission
scheme needs a larger buffer space. In NoC, go-back-N retransmission is
widely used (Ali et al. 2007b; Bertozzi and Benini 2004; Pullini et al. 2005).
7.4.2.3 Error Correction
The error detection and retransmission scheme mentioned above suffers
from higher area requirement as it requires retransmission buffers.
Retransmission will give rise to multiple communications over the same
link, and hence ultimately it will not be energy efficient. Moreover, the performance of the system is degraded due to retransmission. In the DSM era,
area and power are the major issues from the design perspective. A forward
error correction (FEC) technique, however, does not have this type of problem. Single error correction (SEC), for example, Hamming code (Lin and Costello
1983), is widely used as FEC in any VLSI design. The TMR method for error
correction in NoC links was proposed by Huang et al. (2008). Error correction is possible if the Hamming distance between any two code words in the
codebook is greater than one. In general, if the minimum Hamming distance
between any two code words is H, then all the (H – 1) errors appearing on
the bus can be detected and (H – 1)/2 errors can be corrected. Error-correcting
code (ECC) is a linear code (Sridhara and Shanbhag 2005).
With increasing complexity, decreasing power supply voltage, and higher
switching speeds, single-error-correcting capabilities will not sufficiently
increase the error resilience of the system implemented in the current and
forthcoming ultra-DSM (UDSM) technologies. Hence, multiple error-correcting
(MEC) codes are necessary to address this issue. MEC schemes add more
redundancy than the SEC ones. The major challenge in applying the existing MEC schemes in high-speed NoCs are the delay of the codec (encoder
and decoder) module. Bose–Chaudhuri–Hocquenghem (BCH) code, Reed–
Solomon (RS) code, Viterbi code, and Turbo code are widely used for MEC in
off-chip communication. The codec delay of these codes is quite large, and
the summation of wire and codec delay will fail to meet the targeted clock
cycle budget in high-speed NoCs. Hybrid techniques (Murali et al. 2005)
provide both error correction and retransmission and allow for more robust
protection of data. Hybrid solutions compensate for the limitations of ECCs.
For example, SEC and double error detection (DED) codes can correct at most
one error, but can detect double-bit errors. Therefore, upon detection of a
double-bit error, the SEC/DED unit may invoke a retransmission mechanism.
In the current and forthcoming UDSM technologies, the burst error is
likely to better capture the nature of many on-chip errors (Benini and
Micheli 2006). In the work of Zimmer and Jantsch (2003), a conventional fault
model for on-chip buses (Hedge and Shanbhag 2000) has been replaced by a
new fault model to support the burst error. It also proposes a burst error correction technique for NoC links. The overall scheme is shown in Figure 7.22.
