218
Network-on-Chip
However, in the switch-to-switch error detection scheme, error detection is
performed at each switch input port and data retransmission occurs between
adjacent switches. There can be two types of switch-to-switch flow control
schemes: parity or CRC at flit level (ssf) and at packet level (ssp). In the ssp error
detection scheme, the transmitting switch adds parity or CRC bits to the packet’s tail flit, whereas in ssf, the transmitting switch adds parity or CRC bits to
each flit. To handle the fault in ACK/NACK signal (credit signal), a triple modular redundancy (TMR) technique is used as shown in Figure 7.19b. Murali et al.
(2005) showed that the average packet latency is higher in end-to-end rather
than in switch-to-switch flow control. In switch-to-switch flow control, ssp has
more latency than ssf. In NoC, ssf flow control technique is widely used.
The error detection and retransmission methodology is discussed in
Sections 7.4.2.2.1 and 7.4.2.2.2.
7.4.2.2.1 Error Detection Methodology
For error detection, appending a parity bit (odd or even) to the data is the most
common example. But using a parity bit, all burst errors cannot be detected.
However, cyclic redundancy codes (CRCs) take advantage of the considerable burst error detection capability provided by cyclic codes. Common CRC
polynomials can detect the following types of errors: (1) all single-bit errors,
(2) all double-bit errors, (3) all odd number of errors, and (4) any burst error
for which the burst length is less than or equal to the polynomial length.
Linear feedback shift register (LFSR) with serial data feed has been used
since the 1960s to implement the CRC algorithm in hardware. For parallel
data transfer, parallelism has been introduced in CRC-generating hardware
(Albertengo and Sisto 1990; Campobello et al. 2003; Joshi et al. 2000; Pei and
Zukowski 1992; Shieh et al. 2001; Sprachmann et al. 2001). To present the parallel CRC architecture, the LFSR is considered as a synchronous finite-state
machine. State S holds the checksum bits. The message is fed to the input I
and the combinatorial network calculates the next state S next from the current state and the new input. The state machine calculates a new checksum
every clock cycle. For example, the LFSR-based hardware of the generator
polynomial g(x) = x 3 + x 2 + 1 is shown in Figure 7.20a and its parallel implementation is shown in Figure 7.20b. A parallel architecture for the LFSR is
able to include more than one bit of the message in a single clock cycle. For
example, a 4-bit input data word ( i i i i
3 2 1 0 ) with the same generator polynomial
is shown in Figure 7.21.
Therefore, for a 32-bit input, 32 XOR networks will be cascaded to form a
combinational network. The problem of this circuit is that it has large critical
path delay. Thus, the operating clock frequency will degrade. To improve the
performance of this circuit, a byte-wise CRC technique is used (Perez 1983).
In the byte-wise CRC technique, 32-bit data are broken into four bytes, and
for each byte, CRC operation is performed. Thus, four ACK signals will be
generated from the receiver. The ANDing of all the ACK signals is the final
ACK and this signal will come back to the sender.
Précédent

- 237/388

Suivant