An Automatic Sound ClassifiCation Framework …
417
it against other state-of-the-art deep learning and SNN-based models. Finally, we
conclude and discuss the computational benefits of the proposed framework as well
as the considerations when using NVM devices to implement synapses and neurons
in this framework.
2 SOM-SNN ASC Framework
In this section, the feedforward SOM-SNN ASC framework is described. As shown
in Fig. 1, we adopt a biologically plausible auditory front-end (using logarithmic
mel-scaled filter bank that resembles the functionality of the human cochlea) to first
extract low-level spectral features. After which, the unsupervised self-organizing
map (SOM) [18] is used to generate an effective and sparse mid-level feature representation. The best-matching units (BMUs) of the SOM are activated over time
and the corresponding spatiotemporal spike patterns are generated, which represent
the characteristics of each sound event. Finally, the Maximum-Margin Tempotron
temporal learning rule [19] is used to train SNN so as to classify the spike patterns
into different sound categories.
...
Self-Organizing Map
Sound Frame
Signal Preprocessing,
STFT, Log etc.
Mel-scaled Filter
Coefficients
Frequency
Mel-scaled Filter Bank
Energy
...
2.89
1.76
1.54
0.36
0.31
...
Ɵme (ms)
SpaƟotemporal
Spike PaƩern
Spiking Neural Network
...
1
...
Sound
Event 1
Sound
Event i
2
3
4
SOM Neuron Index
Fig. 1 The details of the proposed SOM-SNN ASC framework. The sound frames are pre-processed
and analyzed using mel-scaled filter banks. Then, the SOM generates discrete BMU activation
sequences which are further converted into spike trains. All such spike trains form a spatiotemporal
spike pattern to be classified by the SNN
Précédent

- 417/439

Suivant