440
M. N. Bojnordi and P. Behnam
network workloads. For example, Dian-Nao [55] introduces a parallel multiply-andaccumulate (MAC) unit to exploit the scope of parallelism in CNN and deep neural
network (DNN). This architecture leverages the concept of tiling and prefetching
to reduce the long latency of data movement between main memory and the MAC
units. In a newer version of this architecture, DaDian-Nao [15] extends the original
design by alleviating the challenges of needing huge-memory bandwidth in the
CNN and DNN workloads. Both architectures suffer from a limit performance for
large-scale workloads with excessive memory bandwidth. Eyeriss [56] introduces a
spatial architecture that maximizes data reuse through feeding inputs and weights
to multiple processing elements with local storage and compute units. The design
relies on a hierarchical memory organization that reduces the cost of data movement
from main memory to the processing element. Similarly, ShiDian-Nao [57] employs
a systolic array to maximize the reuse of input data and intermediate results in
computing convolution.
Another important class of energy-efficient accelerators focuses on mapping the
fundamental operations of the machine learning tasks onto analog functional units
inside memory. As a result, these accelerators are able to gain significant speed
and energy-efficiency over fully digital architectures. These accelerators that are
often called analog neuromorphic accelerators rely on leveraging a connectionist
model inspired by the human brain [58] to design the physical structure of
memory and processing elements. Sheri et al. propose a spiking neural network
based on memristive synapses to implement a single-step contrastive divergence
algorithm for machine learning tasks [59]. Each synapse of the proposed design
comprises two memristive elements representing limited-precision positive weights.
Prezioso et al. report the fabrication of a memristive single-level perceptron system
that takes ten inputs to produce three outputs. The circuit is used to classify a
3 × 3 black-and-white image [60]. The memristive Boltzmann Machine [61] and InSitu Analog Arithmetic in Crossbars (ISAAC) [62] propose novel memory-centric
accelerators that perform binary and multibit dot product operations within the
emerging memory arrays. These accelerators propose to eliminate the need for data
movement between memory and computational units. As a result, the memristive
accelerators gain significant performance and energy efficiency. Similarly, PRIME
[63] is proposed as a software/hardware platform for computing matrix-vector
multiplications in resistive arrays for neural networks.
9.2.5 Approximate Computing
Ever since the power consumption was identified as one of the fundamental limitations for the microprocessors’ performance, researchers have examined various
techniques for making computer systems more energy-efficient. Important examples
of such techniques have been considered in IoT systems that include low-power
VLSI circuits for dynamic power and thermal management, dynamic voltage
scaling, multiple-threshold voltage design, and energy harvesting. These techniques
Précédent

- 444/647

Suivant