9 Emerging Hardware Technologies for IoT Data Processing
441
have become even more effective with the help of approximate computing that aims
at balancing accuracy, area, delay, and power consumption based on the user’s
computational needs. Most IoT applications include multimedia processing that
may largely tolerate computational errors without incurring a noticeable quality
loss in output. Therefore, approximate computing is a natural fit for designing
efficient IoT systems. Numerous circuits, architectural mechanisms, and design
methodologies have been proposed in the literature that prove significant power
and performance gains are attainable through applying approximate computing to
different components of IoT systems [64–66]. Recent work on designing imprecise
adders for IoT systems indicates that further performance gains and higher energyefficiencies are attainable through incorporating design techniques that efficiently
explore the design space of approximate units [67].
Designing the most efficient approximate circuit requires specializing both
hardware and software for a given set of design objectives. For example, one can
improve energy efficiency through hardware and software kernels that reduce the
precision of computation [67–69] or reduce the power and energy consumption
through lowering the supply voltage of the existing circuits [70]. The key to a
successful approximate computation is to accurately identify which parts of the
design are error-resilient. This may be done by software through kernel annotations
and compiler techniques [71, 72] or a dedicated approximate data types [73].
Finally, a mechanism is often required to evaluate the quality of result and decide
when to perform an approximation [74, 75].
Machine learning applications seemed largely amenable to approximate computing because of their massive stochastic computation load. Recent work [76]
exploits the inherent redundancy of data and computation within deep CNN layers
and applies a linear compression (singular value decomposition) technique on a
pretrained model to speed up convolution operation during the inference tasks.
Han et al. [77] exploit the sparsity of network parameters via pruning techniques
to reduce the number of redundant weights. The parameters are represented in
a compressed sparse row (CSR) format to increase the storage efficiency. The
architecture is then extended to deep-compression [78] based on quantizing weights
and applying the Huffman encoding to reduce the memory footprint significantly.
Later, a hardware accelerator, called EIE [50], is designed for the compressed
network that achieves substantial speedups and energy savings over prior work due
to executing a set of sequential operations on the compressed data. It is proven
that high precision weights are not important to achieving high accuracy in deep
neural network. Gong et al. [79] propose to quantize the weights of the fully
connected (FC) layer using vector quantization technique at the expense of 1%
accuracy loss. Binarized neural network (BNN) [80] and XNOR-Net [81] extend
the binarization further by binarizing both input and weights that gains significant
reduction in memory footprint and execution time. Tang et al. [82] propose a
resistive full-fledged BNN accelerator that employs binarized hardware for all the
CNN layers. Moreover, Qiu et al. [83] propose an FPGA platform that accelerates
the convolutional neural network for an embedded system.
Précédent

- 445/647

Suivant