RRAM-Based Neuromorphic Computing Systems
403
34.8 μs needed without write-verify scheme. However, the network without writeverify scheme requires 48 more iterations to converge compared to the one with the
write-verify scheme, resulting in more energy consumption during the weight update
phase (197.98–61.16 nJ). This shows that not only the write-verify scheme helped to
mitigate the RRAM variability issue, it also reduced the overall power consumption
during the weight update phase.
The second demonstration was done on cation-based Ta/HfO 2 /Pt RRAM device
grown on 2 μm technology node transistor. Two-layer perceptron was constructed
with 128 × 64 1T1R synapse array with 64 input, 54 hidden, and 10 output neurons.
The network was trained on the handwritten digits from MNIST dataset. The synaptic
weight programming was done by non-identical synchronized pulse scheme to bias
the transistor gate and the top or bottom electrode (TE or BE) of the RRAM cell.
The potentiation was obtained by constant pulse amplitude of 500 μs on the TE with
increasing gate voltage to allow different current flowing through the device. On
the other hand, the depression was performed by resetting the device to the lowest
conductance, by applying 5 μs synchronized pulse to the BE and the gate voltage,
and then implement the conductance increase scheme used during the potentiation
with decreasing transistor gate voltage. This improved the linearity and symmetry of
the device weight update significantly with excellent cycle-to-cycle and device-todevice uniformity [78, 79]. Thus, the network has a promising potential to accommodate in situ training with various learning algorithm. The demonstrated network
was trained by stochastic gradient descent (SGD) to execute the classification task.
The network is initialized through inference by softmax function to obtain the logprobability of each output label before updating the weight within each layer for
every new training dataset. The network underwent 1600 training cycles on 80,000
images from the database. It successfully achieved 91.71% classification accuracy
from 10,000 images.
5 Inference on RRAM Based Neuromorphic Hardware:
From Deep Neural Network (DNN) to Spike-Based
Domain Architecture
a. DNN and SNN
Deep learning has made significant progress in recent years so much so that it has
even outperformed humans in certain tasks, for instance, AlphaGo computer program
managed to defeat the human GO world champion. DNN or deep CNN has achieved
state-of-the-art accuracy in many image classifications or pattern recognition tasks
such as handwritten digit recognition [80], several other datasets such as CIFAR
[81], and ImageNet [82]. However, these networks typically need large amount of
labelled training data; ImageNet has over 1 million labelled images for training.
A conventional CNN is shown in Fig. 9. It comprises of mainly three blocks:
the first block is made up of convolution layers, the second of fully connected
403
34.8 μs needed without write-verify scheme. However, the network without writeverify scheme requires 48 more iterations to converge compared to the one with the
write-verify scheme, resulting in more energy consumption during the weight update
phase (197.98–61.16 nJ). This shows that not only the write-verify scheme helped to
mitigate the RRAM variability issue, it also reduced the overall power consumption
during the weight update phase.
The second demonstration was done on cation-based Ta/HfO 2 /Pt RRAM device
grown on 2 μm technology node transistor. Two-layer perceptron was constructed
with 128 × 64 1T1R synapse array with 64 input, 54 hidden, and 10 output neurons.
The network was trained on the handwritten digits from MNIST dataset. The synaptic
weight programming was done by non-identical synchronized pulse scheme to bias
the transistor gate and the top or bottom electrode (TE or BE) of the RRAM cell.
The potentiation was obtained by constant pulse amplitude of 500 μs on the TE with
increasing gate voltage to allow different current flowing through the device. On
the other hand, the depression was performed by resetting the device to the lowest
conductance, by applying 5 μs synchronized pulse to the BE and the gate voltage,
and then implement the conductance increase scheme used during the potentiation
with decreasing transistor gate voltage. This improved the linearity and symmetry of
the device weight update significantly with excellent cycle-to-cycle and device-todevice uniformity [78, 79]. Thus, the network has a promising potential to accommodate in situ training with various learning algorithm. The demonstrated network
was trained by stochastic gradient descent (SGD) to execute the classification task.
The network is initialized through inference by softmax function to obtain the logprobability of each output label before updating the weight within each layer for
every new training dataset. The network underwent 1600 training cycles on 80,000
images from the database. It successfully achieved 91.71% classification accuracy
from 10,000 images.
5 Inference on RRAM Based Neuromorphic Hardware:
From Deep Neural Network (DNN) to Spike-Based
Domain Architecture
a. DNN and SNN
Deep learning has made significant progress in recent years so much so that it has
even outperformed humans in certain tasks, for instance, AlphaGo computer program
managed to defeat the human GO world champion. DNN or deep CNN has achieved
state-of-the-art accuracy in many image classifications or pattern recognition tasks
such as handwritten digit recognition [80], several other datasets such as CIFAR
[81], and ImageNet [82]. However, these networks typically need large amount of
labelled training data; ImageNet has over 1 million labelled images for training.
A conventional CNN is shown in Fig. 9. It comprises of mainly three blocks:
the first block is made up of convolution layers, the second of fully connected
