An Automatic Sound ClassifiCation Framework …
427
LSF-SNN [26] and LTF-SNN [27] classify the sound samples by first detecting
the spectral features in the power spectrogram, and then encoding these features
into a spatiotemporal spike pattern for classification by a SNN classifier. In our
framework, the SOM is used to learn the key features embedded in the acoustic
signals in an unsupervised manner, which is more biologically plausible. Neurons
in the SOM become selective to specific spectral features after training, and these
features learned by the SOM are more discriminative as shown by the superior SOMSNN classification accuracy compared with the LSF-SNN and LTF-SNN models.
3.2.2 TIDIGITS Dataset
As shown in Table 2, it is encouraging to note that the SOM-SNN framework achieves
an accuracy of 97.40%, outperforming all other bio-inspired systems on the TIDIGITS dataset. In [42–44], novel systems are designed to work with spike streams
generated directly from the AER silicon cochlea sensor. This event-driven auditory
front-end generates spike streams asynchronously from 64 bandpass filters spanning
over the audible range of the human cochlea. Anumula et al. [43] provide a comprehensive overview of the asynchronous and synchronous features generated from
these raw spike streams, once again highlighting the significant role of discriminative
feature representation in speech recognition tasks.
Tavanaei et al. [45, 46] proposes two biologically plausible feature extractors constructed from SNNs trained using the unsupervised spike-timing-dependent plasticity
(STDP) learning rule. The neuronal activations in the feature extraction layer are then
transformed into a real-valued feature vector and used to train a traditional classifier,
such as the HMM or SVM models. In our work, the features are extracted using the
SOM and then used to train a biologically plausible SNN classifier. These differTable 2 Comparison of the classification accuracy of the proposed SOM-SNN framework against
other baseline frameworks on the TIDIGITS dataset
Model
Accuracy (%)
Single-layer SNN and SVM [45] a
91.00
Spiking CNN and HMM [46] a
96.00
AER Silicon Cochlea and SVM [43] b
95.58
AER Silicon Cochlea and Deep RNN [44] b
96.10
AER Silicon Cochlea and Phased LSTM [42] b 91.25
Liquid State Machine [47] c
92.30
MFCC and GRU RNN [42] c
97.90
SOM and SNN (this work) c
97.40
a Evaluate on the Aurora dataset which was developed from the TIDIGITS dataset.
b The data was collected by playing the audio files from the TIDIGITs dataset to the AER Silicon
Cochlea Sensor.
c Evaluate on the TIDIGITS dataset
427
LSF-SNN [26] and LTF-SNN [27] classify the sound samples by first detecting
the spectral features in the power spectrogram, and then encoding these features
into a spatiotemporal spike pattern for classification by a SNN classifier. In our
framework, the SOM is used to learn the key features embedded in the acoustic
signals in an unsupervised manner, which is more biologically plausible. Neurons
in the SOM become selective to specific spectral features after training, and these
features learned by the SOM are more discriminative as shown by the superior SOMSNN classification accuracy compared with the LSF-SNN and LTF-SNN models.
3.2.2 TIDIGITS Dataset
As shown in Table 2, it is encouraging to note that the SOM-SNN framework achieves
an accuracy of 97.40%, outperforming all other bio-inspired systems on the TIDIGITS dataset. In [42–44], novel systems are designed to work with spike streams
generated directly from the AER silicon cochlea sensor. This event-driven auditory
front-end generates spike streams asynchronously from 64 bandpass filters spanning
over the audible range of the human cochlea. Anumula et al. [43] provide a comprehensive overview of the asynchronous and synchronous features generated from
these raw spike streams, once again highlighting the significant role of discriminative
feature representation in speech recognition tasks.
Tavanaei et al. [45, 46] proposes two biologically plausible feature extractors constructed from SNNs trained using the unsupervised spike-timing-dependent plasticity
(STDP) learning rule. The neuronal activations in the feature extraction layer are then
transformed into a real-valued feature vector and used to train a traditional classifier,
such as the HMM or SVM models. In our work, the features are extracted using the
SOM and then used to train a biologically plausible SNN classifier. These differTable 2 Comparison of the classification accuracy of the proposed SOM-SNN framework against
other baseline frameworks on the TIDIGITS dataset
Model
Accuracy (%)
Single-layer SNN and SVM [45] a
91.00
Spiking CNN and HMM [46] a
96.00
AER Silicon Cochlea and SVM [43] b
95.58
AER Silicon Cochlea and Deep RNN [44] b
96.10
AER Silicon Cochlea and Phased LSTM [42] b 91.25
Liquid State Machine [47] c
92.30
MFCC and GRU RNN [42] c
97.90
SOM and SNN (this work) c
97.40
a Evaluate on the Aurora dataset which was developed from the TIDIGITS dataset.
b The data was collected by playing the audio files from the TIDIGITs dataset to the AER Silicon
Cochlea Sensor.
c Evaluate on the TIDIGITS dataset
