An Automatic Sound ClassifiCation Framework …
423
V i (t) as described in Equation 9 , the neurons are encouraged to respond with the
desired spiking activities. This strategy helps to prevent overfitting and improves
classification accuracy.
2.4 Multi-condition Training
Although state-of-the-art deep learning based ASC models perform reasonably well
under the noise-free condition, it remains a challenging task for these models to
recognize sound robustly in noisy real-world environments. To address this challenge,
we investigated training the proposed SOM-SNN model with both clean and noisy
sound data, as per the multi-condition training strategy.
The motivation for such an approach is that with training samples collected from
different noisy backgrounds, the trained model will be encouraged to identify the
most discriminative features and become more robust to noise. This methodology
has been proven to be effective for Deep Neural Network (DNN) and SVM models
under the high noise condition, with some trade-off in performance for clean sound
data [30]. Here, we investigate its generalizability to SNN-based temporal classifiers
under noisy environments.
3 Experimental Results
Here, we first introduce two standard benchmark datasets used to evaluate the classification accuracies of the proposed SOM-SNN framework, which are made up
of environmental sounds and human speech. After which, we describe the experiments conducted on the RWCP dataset to evaluate model performance pertaining
to the effectiveness of feature representation using the SOM, early decision making
capability and noise robustness of the classifier.
3.1 Training and Evaluation Setup
3.1.1 Evaluation Datasets
The Real World Computing Partnership (RWCP) [31] sound scene dataset was
recorded in a real acoustic environment at a sampling rate of 16 kHz. For a fair
comparison with other SNN-based systems [26, 27], we used the same 10 sound
event classes from the dataset: ‘cymbals’, ‘horn’, ‘phone4’, ‘bells5’, ‘kara’, ‘bottle1’, ‘buzzer’, ‘metal15’, ‘whistle1’, ‘ring’. The sound clips were recorded as isolated samples with duration of 0.5s to 3s at high SNR. There are also short lead-in
and lead-out silent intervals in the sound clips. We randomly selected 40 sound clips
Précédent

- 423/439

Suivant