9 Emerging Hardware Technologies for IoT Data Processing
445
Processor Core
Cache
DMA
Controller
Volatile Memory
(e.g., SRAM,
DRAM)
MB-CNN
System Bus
Input Data
Algorithms
and Models
B-CNN Workload
Mobile IoT Device
Fig. 9.9 Overview of an IoT system architecture using the MB-CNN memory-centric accelerator
9.4.2 The MB-CNN Architecture
Designing memory-centric accelerators for deep learning workloads in mobile IoT
devices is challenging because of the increasing demand for more computational
capabilities by the emerging applications. On the other hand, the hardware of
IoT devices is significantly constrained by the stringent cost and power requirements. The memristive binary convolutional neural network (MB-CNN) is a
memory-centric accelerator that addresses this challenge by enabling in situ binary
convolution within the resistive crosspoint memory arrays. The key idea is to exploit
the computational capabilities of the resistive crosspoints for performing the key
operations of the XNOR-Net—i.e., binary XNOR and bit-count. MB-CNN may be
used as a nonvolatile memory system that serves ordinary read and write requests.
Furthermore, the structure of memristive arrays with an additional control logic
allows the framework to perform XNOR convolution at low energy and performance
costs. Figure 9.9 shows how MB-CNN may be employed in a mobile IoT system.
The computational platform includes a microprocessor that executes the application
programs on binary CNN models. A volatile memory module is employed to store
the input data prior to execution. The system is complemented with an MB-CNN
module that accelerates the XNOR convolution and stores the network parameters.
The MB-CNN module is connected to the IoT system via an LPDDR3 standard
memory bus [99]. All the network parameters, such as edge weights, are stored to
the RRAM crosspoint arrays. These parameters are then read to complete multiple
inference tasks. Prior to an XNOR convolution task, a direct memory access (DMA)
controller is used to transfer data from DRAM to the MB-CNN module. For each
layer of the neural network, a set of XNOR convolutions is computed and the
intermediate results are reused for the next layer. At the end of this process, the
software program is responsible to collect the final results from the MB-CNN
module. All the communication between the processor cores and the MB-CNN
module is carried out through LPDDR3 commands.
Précédent

- 449/647

Suivant