9 Emerging Hardware Technologies for IoT Data Processing
453
0
0.2
0.4
0.6
0.8
1
1.2
1.4
1.6
1.8
2
0
0.2
0.4
0.6
0.8
1
Relative Execution Time
Relative System Energy
CPU
GPU
PIM
MB-CNN
S
D
Q
S
D
Q
Q
S
D
Q
S
D
Fig. 9.17 Relative execution time and system energy of MB-CNN compared to the baseline
architectures with single (S)-, dual (D)-, and quad (Q)-core processors
a better performance than the GPU-based systems, MB-CNN outperforms the PIMlike ASICs by 1.29×, 1.58×, and 2× in the single-, dual-, and quad-core systems,
respectively. Notice that MB-CNN benefits from in situ computation within memory
arrays that enables massive parallelism and eliminates unnecessary data movement.
Moreover, MB-CNN achieves average energy savings of 4×, 3.83×, and 3.64×
over the CPU baselines with single-, dual-, and quad-core processors. The GPUbased systems consume more chip area and power to improve performance.
However, they can only achieve a better energy saving that the PIM accelerator with
a quad-core processor. This is mainly because of the reduced leakage energy in the
GPU as the overall execution time on the processor decreased. MB-CNN achieves
better energy savings over the PIM- and GPU-based accelerators by respective
2.46×, 2.61×, and 2.63× for the single-, dual-, and quad-core systems, respectively.
9.5 In Situ Data Clustering for IoT Servers
This section presents the memristive in-situ clustering (MISC) architecture as
another example of in situ accelerators. The MISC architecture is designed to
perform energy-efficient and fast data clustering within memristive arrays. The
memory arrays are specifically re-structured to support the basic operations of
MISC. Moreover, algorithmic techniques and special design strategies are considered to enable large-scale data clustering on MISC.
453
0
0.2
0.4
0.6
0.8
1
1.2
1.4
1.6
1.8
2
0
0.2
0.4
0.6
0.8
1
Relative Execution Time
Relative System Energy
CPU
GPU
PIM
MB-CNN
S
D
Q
S
D
Q
Q
S
D
Q
S
D
Fig. 9.17 Relative execution time and system energy of MB-CNN compared to the baseline
architectures with single (S)-, dual (D)-, and quad (Q)-core processors
a better performance than the GPU-based systems, MB-CNN outperforms the PIMlike ASICs by 1.29×, 1.58×, and 2× in the single-, dual-, and quad-core systems,
respectively. Notice that MB-CNN benefits from in situ computation within memory
arrays that enables massive parallelism and eliminates unnecessary data movement.
Moreover, MB-CNN achieves average energy savings of 4×, 3.83×, and 3.64×
over the CPU baselines with single-, dual-, and quad-core processors. The GPUbased systems consume more chip area and power to improve performance.
However, they can only achieve a better energy saving that the PIM accelerator with
a quad-core processor. This is mainly because of the reduced leakage energy in the
GPU as the overall execution time on the processor decreased. MB-CNN achieves
better energy savings over the PIM- and GPU-based accelerators by respective
2.46×, 2.61×, and 2.63× for the single-, dual-, and quad-core systems, respectively.
9.5 In Situ Data Clustering for IoT Servers
This section presents the memristive in-situ clustering (MISC) architecture as
another example of in situ accelerators. The MISC architecture is designed to
perform energy-efficient and fast data clustering within memristive arrays. The
memory arrays are specifically re-structured to support the basic operations of
MISC. Moreover, algorithmic techniques and special design strategies are considered to enable large-scale data clustering on MISC.
