452
M. N. Bojnordi and P. Behnam
to the crosspoint per XNOR convolution. The inputs are then distributed among the
four arrays. Due to using a 5-bit sensor, only 32 rows (cell segment) of the arrays
are driven by the input data. All the other rows remain inactive and do not contribute
to the in situ computation.
As mentioned in Sect. 4.4.2, the 5-bit partial results from the arrays are given
to the serial adders of the reduction tree to compute the final outputs {z 0 , z 1 ... z 35 },
where zn is the result of a convolution between input I and kernel Kn. An internal
control mechanism is employed to switch a new set of 32 cells into each sensor every
five cycles. The switching operation happens in a pipelined fashion inside the bank
to produce one output element (zn) every seven cycles. Ultimately the first element
of all output feature becomes available at local output buffer after 7 × 128 cycles.
The chip controller feeds the next set of inputs to the memristive arrays
for convolution once the current subset is reused by all convolution filters.
Therefore, producing the complete output of a convolution layer takes
7 × 128 × 6 × 6 = 32,256 cycles. Notice that more concurrent computation is
possible through increasing the number of sense amplifiers per memristive arrays
and by replicating the same kernel parameters across multiple banks. However, that
improved performance requires more chip area and power consumption.
9.4.5 Potentials of the MB-CNN Accelerator
The MB-CNN accelerator can be integrated in mobile systems with single- or
multicore processors. This section examines the energy and performance potentials
of the accelerator used by single- and multicore processors that realize the MIPS64
instruction set architecture (ISA). For better evaluations, a GPU-based solution and
an ASIC accelerator for processing-in-memory (PIM) are considered as the baseline
systems for comparisons. The GPU solution is based on the Nvidia Tegra X1 lowpower system with 256 processing cores [53]. The low power GPU is mainly used
to implement the floating-point convolutions in the first and last layers of an endto-end inference task. The PIM ASIC solution integrates additional gates near the
RRAM arrays to implement the XNOR trees and bit-counters that fetch data from
the arrays and compute the XNOR convolution. Like the MB-CNN hardware, the
outcome of each XNOR convolution is transferred to software for completing the
layer computation. Moreover, the PIM baseline is optimized so that it occupies
the same area as that of MB-CNN; however, it does not support in situ XNOR
computation.
Figure 9.17 shows the relative execution time and system energy of the XNORNet inference across various system configurations, namely CPU, GPU, PIM, and
MB-CNN. For each configuration, three design points with respect to the number
of processor cores—i.e., single (S), dual (D), and quad (Q)—are considered. MBCNN outperforms all of the baseline systems. As compared with CPU, MB-CNN
achieves 4.17×, 4.25×, and 3.71× performance improvements for the single-,
dual-, and quad-core systems, respectively. Although the PIM accelerator achieve
M. N. Bojnordi and P. Behnam
to the crosspoint per XNOR convolution. The inputs are then distributed among the
four arrays. Due to using a 5-bit sensor, only 32 rows (cell segment) of the arrays
are driven by the input data. All the other rows remain inactive and do not contribute
to the in situ computation.
As mentioned in Sect. 4.4.2, the 5-bit partial results from the arrays are given
to the serial adders of the reduction tree to compute the final outputs {z 0 , z 1 ... z 35 },
where zn is the result of a convolution between input I and kernel Kn. An internal
control mechanism is employed to switch a new set of 32 cells into each sensor every
five cycles. The switching operation happens in a pipelined fashion inside the bank
to produce one output element (zn) every seven cycles. Ultimately the first element
of all output feature becomes available at local output buffer after 7 × 128 cycles.
The chip controller feeds the next set of inputs to the memristive arrays
for convolution once the current subset is reused by all convolution filters.
Therefore, producing the complete output of a convolution layer takes
7 × 128 × 6 × 6 = 32,256 cycles. Notice that more concurrent computation is
possible through increasing the number of sense amplifiers per memristive arrays
and by replicating the same kernel parameters across multiple banks. However, that
improved performance requires more chip area and power consumption.
9.4.5 Potentials of the MB-CNN Accelerator
The MB-CNN accelerator can be integrated in mobile systems with single- or
multicore processors. This section examines the energy and performance potentials
of the accelerator used by single- and multicore processors that realize the MIPS64
instruction set architecture (ISA). For better evaluations, a GPU-based solution and
an ASIC accelerator for processing-in-memory (PIM) are considered as the baseline
systems for comparisons. The GPU solution is based on the Nvidia Tegra X1 lowpower system with 256 processing cores [53]. The low power GPU is mainly used
to implement the floating-point convolutions in the first and last layers of an endto-end inference task. The PIM ASIC solution integrates additional gates near the
RRAM arrays to implement the XNOR trees and bit-counters that fetch data from
the arrays and compute the XNOR convolution. Like the MB-CNN hardware, the
outcome of each XNOR convolution is transferred to software for completing the
layer computation. Moreover, the PIM baseline is optimized so that it occupies
the same area as that of MB-CNN; however, it does not support in situ XNOR
computation.
Figure 9.17 shows the relative execution time and system energy of the XNORNet inference across various system configurations, namely CPU, GPU, PIM, and
MB-CNN. For each configuration, three design points with respect to the number
of processor cores—i.e., single (S), dual (D), and quad (Q)—are considered. MBCNN outperforms all of the baseline systems. As compared with CPU, MB-CNN
achieves 4.17×, 4.25×, and 3.71× performance improvements for the single-,
dual-, and quad-core systems, respectively. Although the PIM accelerator achieve
