416
J. Wu et al.
Driven by large-scale deep learning models, abundant training data and powerful parallel graphics processing units (GPUs), deep learning has made remarkable
progress in automatic sound classification. Despite compelling performance has been
demonstrated with these deep learning models [5], the required high-performance
computing, which usually comes along with high power consumption, prevent the
deployment of these models on the pervasive mobile and wearable devices. Furthermore, the performance of these deep learning based system degrades significantly
with the increasing amount of environmental noise.
Humans perform efficiently and robustly in various auditory perception tasks,
whereby spectral contents of the acoustic signal are encoded asynchronously using
sparse and highly-parallel spiking impulses. Remarkably, despite the fact that spiking
impulses in biological neural systems transmit at several orders of magnitude slower
than the signals in modern transistors, humans are able to analyze complex audio
scenes effortlessly with much lower energy consumption [6]. Moreover, humans
learn to distinguish sounds with only sparse supervision, with occasional labeled
data as in zero-shot or one-shot learning [7, 8].
The event-based computation, as has been observed in the human brain and sensory systems, relies on asynchronous and highly parallel spiking events to efficiently
represent information. In contrast to traditional frame-based machine vision and auditory systems, event-based biological neural systems represent and process information in a much more energy efficient manner with energy consumed only during generation and transmission of spikes. Notably, neuromorphic computing has emerged
with the vision to mimic such event-based biological neural systems. Neuromorphic
computing leverages on low-power, densely-connected parallel computing units to
perform complex perceptual and cognitive tasks; the inherent collocating memory
and processing effectively address the problem of low bandwidth between the CPU
and memory (i.e., von Neumann bottleneck) [9].
The emerging dense crossbar array of non-volatile memory (NVM) devices have
been recognized as a promising approach to emulate such distributed, massivelyparallel and densely connected neuromorphic computing systems in hardware[10–
12]. The compact, low power NVM devices with continuous, near-linear conductance dynamic range are attractive for implementing synapses and neurons, instances
include Phase Change Memory (PCM) and Resistive RAM (RRAM) etc. When integrated with event-based sensors, such as the DVS [13], DAVIS [14] and DAS [15],
such event-based neuromorphic computing systems are attractive for real-time, adaptive and energy efficient applications [16, 17].
In this chapter, we introduce a novel neuromorphic automatic sound classification
framework based on the spiking neural network (SNN). Moreover, we explain how to
implement this framework with emerging NVM devices, organically integrating the
compelling algorithmic power with efficient hardware for real-world applications.
The rest of this chapter is organized as follows: we first describe the system architecture of the proposed ASC framework, which uses unsupervised self-organizing
map (SOM) for feature representation and SNN for temporal classification, namely
SOM-SNN. Subsequently, we present the experiments designed to evaluate the classification performance and robustness to noise of the proposed framework; compare
Précédent

- 416/439

Suivant