9 Emerging Hardware Technologies for IoT Data Processing
439
Artificial Intelligence:techniques
that enable machines to mimic
cognitive functions like humans.
Deep Learning: a subset of
machine learning relying on
multi-layer structures.
Machine Learning: statistical
methods for performing tasks
based on inference.
Fig. 9.5 Deep machine learning as a significant fraction of artificial intelligence
9.2.4 Machine Learning Accelerators in the IoT Era
The core application of most IoT devices is to detect different human behaviors,
sense the ambient contexts, and produce the appropriate reaction. Machine learning
has emerged as a key technique to enable these applications through extracting
sensor data, identifying meaningful context, and performing intelligent tasks for
face detection [42], image classification [43], and speech recognition [44]. As
shown in Fig. 9.5, machine learning is a significant fraction of artificial intelligence
(AI). Deep machine learning techniques, such as convolutional neural network
(CNN), have emerged as the most successful class of machine learning that rely
on multilayer neural networks for computation.
Due to the increasing demand of computation, memory, and energy consumption
of machine learning applications, engineers and researchers have considered designing efficient ways to accelerate machine learning workloads. Software libraries
have been proposed to accelerate deep learning tasks such as speech recognition
(speech-to-text and speech-to-command) and computer vision (face detection and
image classification) on low-power mobile GPUs [45, 46]. Recent research work on
wear bench application shows that out-of-order processor cores may be adequate
to achieve a high performance for deep learning workloads on wearable IoT
devices [47]. To maximize the forward progress of IoT applications in unstable
and intermittent power supply environment, nonvolatile processor architectures have
recently been considered in the literature [48].
As an important class of deep learning, CNN has proven successful in image
classification and face recognition [43, 49]. A typical deep convolutional neural
network may require millions of parameters to be learned and stored during
the training phase and to be retrieved and used for inference tasks. In addition
to the stringent memory and storage requirements of the IoT nodes, accessing
these parameters by various layers of the neural network necessitates consuming
significant amounts of energy and time. Interestingly, high-precision parameters
are not important to gain high accuracy in the outcome of a neural network;
as a result, numerous techniques have been proposed in the literature that focus
on trading the computation precision for achieving better energy efficiency and
performance [50–54]. In addition, high-performance and energy-efficient hardware
accelerators have been considered for computation and memory-intensive neural
439
Artificial Intelligence:techniques
that enable machines to mimic
cognitive functions like humans.
Deep Learning: a subset of
machine learning relying on
multi-layer structures.
Machine Learning: statistical
methods for performing tasks
based on inference.
Fig. 9.5 Deep machine learning as a significant fraction of artificial intelligence
9.2.4 Machine Learning Accelerators in the IoT Era
The core application of most IoT devices is to detect different human behaviors,
sense the ambient contexts, and produce the appropriate reaction. Machine learning
has emerged as a key technique to enable these applications through extracting
sensor data, identifying meaningful context, and performing intelligent tasks for
face detection [42], image classification [43], and speech recognition [44]. As
shown in Fig. 9.5, machine learning is a significant fraction of artificial intelligence
(AI). Deep machine learning techniques, such as convolutional neural network
(CNN), have emerged as the most successful class of machine learning that rely
on multilayer neural networks for computation.
Due to the increasing demand of computation, memory, and energy consumption
of machine learning applications, engineers and researchers have considered designing efficient ways to accelerate machine learning workloads. Software libraries
have been proposed to accelerate deep learning tasks such as speech recognition
(speech-to-text and speech-to-command) and computer vision (face detection and
image classification) on low-power mobile GPUs [45, 46]. Recent research work on
wear bench application shows that out-of-order processor cores may be adequate
to achieve a high performance for deep learning workloads on wearable IoT
devices [47]. To maximize the forward progress of IoT applications in unstable
and intermittent power supply environment, nonvolatile processor architectures have
recently been considered in the literature [48].
As an important class of deep learning, CNN has proven successful in image
classification and face recognition [43, 49]. A typical deep convolutional neural
network may require millions of parameters to be learned and stored during
the training phase and to be retrieved and used for inference tasks. In addition
to the stringent memory and storage requirements of the IoT nodes, accessing
these parameters by various layers of the neural network necessitates consuming
significant amounts of energy and time. Interestingly, high-precision parameters
are not important to gain high accuracy in the outcome of a neural network;
as a result, numerous techniques have been proposed in the literature that focus
on trading the computation precision for achieving better energy efficiency and
performance [50–54]. In addition, high-performance and energy-efficient hardware
accelerators have been considered for computation and memory-intensive neural
