30
Network-on-Chip
certain flit faces a busy channel, subsequent flits also have to wait at their
current locations.
In the absence of contention, VCT and wormhole switching have the same
latency. Otherwise, VCT has lower latency and higher acceptance rate compared to wormhole switching (Banerjee et al. 2004). However, VCT requires
larger silicon area and consumes higher energy due to larger buffer size.
Therefore, a trade-off is necessary between energy consumption, area, and
performance. Wormhole switching is preferable for large packet size, whereas
VCT switching is a better choice for short packets. When wormhole switching is employed, the header flit gets blocked if the output channel is already
assigned to another packet. This problem is known as head-of-line (HoL)
blocking. This affects the overall system performance. To mitigate this effect,
Dally (1992) proposed the usage of several virtual channels (VCs) within each
physical channel. When a particular packet is blocked, VC allows other packets to use the link that would otherwise be left idle. Usage of VC improves
the overall performance at the cost of increased energy consumption and
silicon area overhead. However, VC cannot eliminate the HoL blocking problem completely. Duato et al. (2003) compared the performance of VCT- and
VC-based wormhole switching, both having an equal buffer capacity. It has
been observed that the VC-based wormhole switching achieves much higher
throughput, whereas the average latency in both the cases is almost identical
before saturation. In NoC design, wormhole routers with limited number
of VCs are preferable. The optimum number of VCs per physical channel is
determined by power–performance trade-off of the overall system. Pande
et al. (2005) reported that the optimum value of VCs per physical channel is
4, beyond which the throughput increment is marginal, whereas the energy
consumption and area overhead increase. To reduce zero-load latency and
router energy in a VC-based network, Kumar et al. (2007) introduced the
express cube structure that allows packets to virtually bypass the intermediate routers along their path in a completely nonspeculative fashion.
2.4 Routing Strategies
Routing strategies determine the traversal path of a packet from the source
to the destination. Depending on the number of destinations of a single
packet, routing algorithm can be classified as unicast and multicast. In unicast
routing, each packet has a single destination, whereas in multicast routing, a
single packet has multiple destinations. For on-chip communication, unicast
routing seems to be a practical approach due to the presence of point-to-point
communication links between various components inside a chip (Agarwal
et al. 2009). Routing techniques can be further classified as source and distributed based on the position at which routing decision takes place. In source
Network-on-Chip
certain flit faces a busy channel, subsequent flits also have to wait at their
current locations.
In the absence of contention, VCT and wormhole switching have the same
latency. Otherwise, VCT has lower latency and higher acceptance rate compared to wormhole switching (Banerjee et al. 2004). However, VCT requires
larger silicon area and consumes higher energy due to larger buffer size.
Therefore, a trade-off is necessary between energy consumption, area, and
performance. Wormhole switching is preferable for large packet size, whereas
VCT switching is a better choice for short packets. When wormhole switching is employed, the header flit gets blocked if the output channel is already
assigned to another packet. This problem is known as head-of-line (HoL)
blocking. This affects the overall system performance. To mitigate this effect,
Dally (1992) proposed the usage of several virtual channels (VCs) within each
physical channel. When a particular packet is blocked, VC allows other packets to use the link that would otherwise be left idle. Usage of VC improves
the overall performance at the cost of increased energy consumption and
silicon area overhead. However, VC cannot eliminate the HoL blocking problem completely. Duato et al. (2003) compared the performance of VCT- and
VC-based wormhole switching, both having an equal buffer capacity. It has
been observed that the VC-based wormhole switching achieves much higher
throughput, whereas the average latency in both the cases is almost identical
before saturation. In NoC design, wormhole routers with limited number
of VCs are preferable. The optimum number of VCs per physical channel is
determined by power–performance trade-off of the overall system. Pande
et al. (2005) reported that the optimum value of VCs per physical channel is
4, beyond which the throughput increment is marginal, whereas the energy
consumption and area overhead increase. To reduce zero-load latency and
router energy in a VC-based network, Kumar et al. (2007) introduced the
express cube structure that allows packets to virtually bypass the intermediate routers along their path in a completely nonspeculative fashion.
2.4 Routing Strategies
Routing strategies determine the traversal path of a packet from the source
to the destination. Depending on the number of destinations of a single
packet, routing algorithm can be classified as unicast and multicast. In unicast
routing, each packet has a single destination, whereas in multicast routing, a
single packet has multiple destinations. For on-chip communication, unicast
routing seems to be a practical approach due to the presence of point-to-point
communication links between various components inside a chip (Agarwal
et al. 2009). Routing techniques can be further classified as source and distributed based on the position at which routing decision takes place. In source
