442
M. N. Bojnordi and P. Behnam
Fig. 9.6 Relative system
performance and energy
consumption of various
architectures
Relative Performance
Relative Energy Consumption
NMP
ISP
ASIC
FPGA
GPU
CPU
9.3 Near-Memory Processing
Most accelerators employ various techniques for in-memory processing that minimize data movement between memory and processor cores to provide significant
energy savings and performance improvements. In-memory processing is an old
concept that has been revisited recently by both industry and academia in the
advent of big data computing and recent advances of technology—e.g., die stacking,
emerging nonvolatile memories, and high-bandwidth memory interfaces. Based
on the location of computation with respect to the memory cells, data-centric
accelerators may be divided into near-memory processing (NMP) [84, 85] and in
situ processing (ISP) [61, 62, 86]. Figure 9.6 illustrates the relative performance
and energy consumption of various design approaches for hardware accelerators
and general purpose processors based on the results from recent work on big data
processing [84, 87], machine learning acceleration [43, 62, 88], and optimization
problems [61, 89]. As shown in the figure, ISP and ASIC NMP provide significantly
better performance and energy savings compared to other techniques. This superior
energy-efficiency is achievable mainly because of exploiting the unprecedented
parallelism at the level of memory arrays and cells while reducing data movement
to a minimal amount through performing digital processing at the periphery of data
arrays or analog computation within memory cells. Examples of these architectures
are (1) computing bitwise Boolean functions, such as NOR, within DRAM arrays
in DRISA [90], (2) utilizing the inherent dot-product capability of memory arrays
to accelerate matrix-vector multiplication in ISAAC [62] and the memristive
Boltzmann machine [61], and (3) performing associative search operations inside
data arrays to realize TCAM-DIMM [91] and AC-DIMM [87]. In the rest of this
section, we will explain the design of two ISP architectures based on memristive
technology for big data processing in IoT systems. First, we introduce an energyefficient memory system capable of accelerating binary neural network tasks in
mobile IoT devices [92]. Then, we examine the architecture of an ISP system for
large-scale data clustering for IoT servers and data centers [93].
Précédent

- 446/647

Suivant