426
J. Wu et al.
Neuronal Noise We also consider the effect of neuronal noise which is known to
exist in the human brain, emulated by spike jittering and deletion. Given that the
human auditory system is highly robust to these noises, it motivates us to investigate
the performance of the proposed framework under such noisy conditions.
For spike jittering, we add Gaussian noise with zero mean and standard deviation
σ to the spike timing t of all input spikes entering the SNN classifier. The amount
of jitter is determined by σ which we sweep from 0.1 T to 0.8 T , where T is the
spike generation period. In addition, we also consider spike deletion, where a certain
fraction of spikes are corrupted by noise and not delivered to the SNN. For both types
of neuronal noise, we trained the model without any noise and then tested it with
jittered (of varying standard deviation σ ) or deleted (of varying ratio) input spike
trains.
3.2 Classification Results
3.2.1 RWCP Dataset
As shown in Table 1, the SOM-SNN model achieved a test accuracy of 99.60%,
which is competitive compared with other deep learning and SNN-based models. As
described in the experimental set-up, the MLP and CNN models are trained using
spectrogram images of fixed dimensions, instead of explicitly modeling the temporal
transition of frames. Despite their high accuracy on this dataset, it may be challenging
to use them for classifying sound samples of long duration; the temporal structures
will be affected inconsistently due to the necessary rescaling of the spectrogram
images [40]. On the other hand, the RNN and LSTM models capture the temporal
transition explicitly. These models are however hard to train for long sound samples
due to the vanishing and exploding gradient problem [41].
Table 1 Comparison of the classification accuracy of the proposed SOM-SNN framework against
other ANNs and SNN-based frameworks on the RWCP dataset. The average results over 10 experimental runs with random weight initialization are reported
Model
Accuracy (%)
MLP
99.45
CNN
99.85
RNN
95.35
LSTM
98.40
LSF-SNN [26]
98.50
LTF-SNN [27]
97.50
SOM-SNN (ReSuMe)
97.00
SOM-SNN (Maximum-Margin Tempotron)
99.60
J. Wu et al.
Neuronal Noise We also consider the effect of neuronal noise which is known to
exist in the human brain, emulated by spike jittering and deletion. Given that the
human auditory system is highly robust to these noises, it motivates us to investigate
the performance of the proposed framework under such noisy conditions.
For spike jittering, we add Gaussian noise with zero mean and standard deviation
σ to the spike timing t of all input spikes entering the SNN classifier. The amount
of jitter is determined by σ which we sweep from 0.1 T to 0.8 T , where T is the
spike generation period. In addition, we also consider spike deletion, where a certain
fraction of spikes are corrupted by noise and not delivered to the SNN. For both types
of neuronal noise, we trained the model without any noise and then tested it with
jittered (of varying standard deviation σ ) or deleted (of varying ratio) input spike
trains.
3.2 Classification Results
3.2.1 RWCP Dataset
As shown in Table 1, the SOM-SNN model achieved a test accuracy of 99.60%,
which is competitive compared with other deep learning and SNN-based models. As
described in the experimental set-up, the MLP and CNN models are trained using
spectrogram images of fixed dimensions, instead of explicitly modeling the temporal
transition of frames. Despite their high accuracy on this dataset, it may be challenging
to use them for classifying sound samples of long duration; the temporal structures
will be affected inconsistently due to the necessary rescaling of the spectrogram
images [40]. On the other hand, the RNN and LSTM models capture the temporal
transition explicitly. These models are however hard to train for long sound samples
due to the vanishing and exploding gradient problem [41].
Table 1 Comparison of the classification accuracy of the proposed SOM-SNN framework against
other ANNs and SNN-based frameworks on the RWCP dataset. The average results over 10 experimental runs with random weight initialization are reported
Model
Accuracy (%)
MLP
99.45
CNN
99.85
RNN
95.35
LSTM
98.40
LSF-SNN [26]
98.50
LTF-SNN [27]
97.50
SOM-SNN (ReSuMe)
97.00
SOM-SNN (Maximum-Margin Tempotron)
99.60
