9 Emerging Hardware Technologies for IoT Data Processing
443
Processor Core
Cache
DMA
Controller
Volatile Memory
(e.g., SRAM,
DRAM)
Non-volatile
Memory (e.g.,
FLASH, RRAM)
System Bus
Network
User IO
Fig. 9.7 Illustrative example of the system architecture of an IoT device
9.4 In Situ Processing for IoT Devices
IoT devices demand critical optimization for stringent ultra-low power requirements, which makes the realization of data-intensive applications, such as a
full-fledged deep learning application, a significant challenge. Figure 9.7 shows
an illustrative example of a generic system architecture for IoT devices and
edge nodes. Depending on the application objectives and the design constraints,
a single- or multicore processor is employed for executing the user programs.
Typically, the memory system consists of both volatile and nonvolatile subsystems
that are interfaced to the processor cores via a system bus. Single- or multiplecache levels as well as a scratchpad memory may be used as fast and temporary
memory. Nonvolatile memory is commonly used to store application programs and
data permanently. Nowadays, FLASH is a widely used technology for building
nonvolatile memory in wearable and mobile devices [94]. The emerging nonvolatile
memory technologies such as FeRAM [95] and RRAM [96] are expected to replace
the conventional FLASH memories in future [97, 98].
9.4.1 Deep Binary Neural Network
Given the strict power and performance requirements of IoT devices, binarized
CNN seems a promising model of deep neural networks in edge computing. In a
binary deep neural network, both the inputs and weights are binarized. For example,
XNOR-Net is a binary neural network that consists of four main stages, namely
Batch Normalizations, Binary Activation, XNOR Convolution, and Pooling. XNORNet was initially based on converting the parameters into either +1 or −1 using a
sign function defined as the following:
x
b
=
+1 x ≥ 0
− 1 x < 0
443
Processor Core
Cache
DMA
Controller
Volatile Memory
(e.g., SRAM,
DRAM)
Non-volatile
Memory (e.g.,
FLASH, RRAM)
System Bus
Network
User IO
Fig. 9.7 Illustrative example of the system architecture of an IoT device
9.4 In Situ Processing for IoT Devices
IoT devices demand critical optimization for stringent ultra-low power requirements, which makes the realization of data-intensive applications, such as a
full-fledged deep learning application, a significant challenge. Figure 9.7 shows
an illustrative example of a generic system architecture for IoT devices and
edge nodes. Depending on the application objectives and the design constraints,
a single- or multicore processor is employed for executing the user programs.
Typically, the memory system consists of both volatile and nonvolatile subsystems
that are interfaced to the processor cores via a system bus. Single- or multiplecache levels as well as a scratchpad memory may be used as fast and temporary
memory. Nonvolatile memory is commonly used to store application programs and
data permanently. Nowadays, FLASH is a widely used technology for building
nonvolatile memory in wearable and mobile devices [94]. The emerging nonvolatile
memory technologies such as FeRAM [95] and RRAM [96] are expected to replace
the conventional FLASH memories in future [97, 98].
9.4.1 Deep Binary Neural Network
Given the strict power and performance requirements of IoT devices, binarized
CNN seems a promising model of deep neural networks in edge computing. In a
binary deep neural network, both the inputs and weights are binarized. For example,
XNOR-Net is a binary neural network that consists of four main stages, namely
Batch Normalizations, Binary Activation, XNOR Convolution, and Pooling. XNORNet was initially based on converting the parameters into either +1 or −1 using a
sign function defined as the following:
x
b
=
+1 x ≥ 0
− 1 x < 0
