9 Emerging Hardware Technologies for IoT Data Processing
461
omitted. This optimization is considered by dedicating a low memristive element in
b i−1 , which is included in the current summation only if P = 1. After completing
the majority vote at the current bit position (b i ), P and I are recomputed for the next
bit position (b i+1 ). Figure 9.24b shows how P and I are recomputed using b i . The
true and complement values of the newly computed majority vote are applied to M
and M. Then, E is connected to Vdd to enable the XNOR part of the cell. The result
of XNORing M and R is produced on the wordline I. The wordline is connected
to a control circuit that detects a 1-to-0 transition and locks it to 0 till the end of
computation. Moreover, the control circuit sets P to 1 only if M is equal to 1.
Updating the Cell Recall that the MISC cell stores the true and complement values
of the data bits; therefore, updating the contents of every cell requires additional
writes. MISC employs a two-phase update mechanism that writes all 1s in the first
phase and then all of the 0s. This process does not incur significant overhead in
a data clustering problem because the dataset is written in the memory once and
is read by the algorithm multiple times. Moreover, the performance and energy
benefits of in situ computing surpass this overhead significantly.
9.5.5.2 Analog Bit Counter and Reduction Network
Theoretically, solving a large-scale data clustering problem with MISC needs
computing the majority vote of a large number of data points stored in a single
memory array. Building large MISC arrays is impractical due to significant sensing
and reliability issues. Instead, MISC stores data points in multiple limited sized
arrays, and only a fraction of the cells within each column is processed using the
analog bit counters. The multibit sensors are similar to those used in the MBCNN arrays. Multiple MISC array computations are performed in parallel to gain
significant performance. Again, a hierarchical merging mechanism is proposed to
compute the majority vote of many data points stored in multiple MISC arrays. An
interconnection network comprising reduction units merges the partial bit-counts
computed per arrays into a single majority bit. The main purpose of the reduction
tree is to merge the partial counts computed by the analog bit counters.
A reconfigurable reduction tree is used inside each bank to interconnect the data
arrays and the chip controller. The tree is capable of selectively merging the partial
counts from the data arrays into a single count value. Figure 9.25 shows the MISC
reduction unit with nine possible ways of reading data from the children arrays A,
B, C, and D. Each MISC reduction unit is configured using a 2-bit mode register
(m). By sharing the modes values among the reduction units of each layer, nine
useful configurations are possible for reducing the partial results in MISC. Similar
to MB-CNN, the nodes are programmed to appropriate operational modes prior to
a computation task. MISC makes it now possible to read the individual arrays that
are used for serving ordinary read requests or to read the sum of values provided
by every two or four adjacent arrays. Such flexibility has been essential to achieve
significant energy efficiency in solving problems that partially occupy the MISC
461
omitted. This optimization is considered by dedicating a low memristive element in
b i−1 , which is included in the current summation only if P = 1. After completing
the majority vote at the current bit position (b i ), P and I are recomputed for the next
bit position (b i+1 ). Figure 9.24b shows how P and I are recomputed using b i . The
true and complement values of the newly computed majority vote are applied to M
and M. Then, E is connected to Vdd to enable the XNOR part of the cell. The result
of XNORing M and R is produced on the wordline I. The wordline is connected
to a control circuit that detects a 1-to-0 transition and locks it to 0 till the end of
computation. Moreover, the control circuit sets P to 1 only if M is equal to 1.
Updating the Cell Recall that the MISC cell stores the true and complement values
of the data bits; therefore, updating the contents of every cell requires additional
writes. MISC employs a two-phase update mechanism that writes all 1s in the first
phase and then all of the 0s. This process does not incur significant overhead in
a data clustering problem because the dataset is written in the memory once and
is read by the algorithm multiple times. Moreover, the performance and energy
benefits of in situ computing surpass this overhead significantly.
9.5.5.2 Analog Bit Counter and Reduction Network
Theoretically, solving a large-scale data clustering problem with MISC needs
computing the majority vote of a large number of data points stored in a single
memory array. Building large MISC arrays is impractical due to significant sensing
and reliability issues. Instead, MISC stores data points in multiple limited sized
arrays, and only a fraction of the cells within each column is processed using the
analog bit counters. The multibit sensors are similar to those used in the MBCNN arrays. Multiple MISC array computations are performed in parallel to gain
significant performance. Again, a hierarchical merging mechanism is proposed to
compute the majority vote of many data points stored in multiple MISC arrays. An
interconnection network comprising reduction units merges the partial bit-counts
computed per arrays into a single majority bit. The main purpose of the reduction
tree is to merge the partial counts computed by the analog bit counters.
A reconfigurable reduction tree is used inside each bank to interconnect the data
arrays and the chip controller. The tree is capable of selectively merging the partial
counts from the data arrays into a single count value. Figure 9.25 shows the MISC
reduction unit with nine possible ways of reading data from the children arrays A,
B, C, and D. Each MISC reduction unit is configured using a 2-bit mode register
(m). By sharing the modes values among the reduction units of each layer, nine
useful configurations are possible for reducing the partial results in MISC. Similar
to MB-CNN, the nodes are programmed to appropriate operational modes prior to
a computation task. MISC makes it now possible to read the individual arrays that
are used for serving ordinary read requests or to read the sum of values provided
by every two or four adjacent arrays. Such flexibility has been essential to achieve
significant energy efficiency in solving problems that partially occupy the MISC
