An Automatic Sound ClassifiCation Framework …
425
3.1.3 Traditional Artificial Neural Networks
To facilitate comparison with other traditional ANN models trained on the RWCP
dataset, we implement four common neural network architectures, namely the MultiLayer Perceptron (MLP) [33], the Convolutional Neural Network (CNN) [34], the
Recurrent Neural Network (RNN) [35] and the Long Short-Term Memory (LSTM)
[36] using the Pytorch library. For a fair comparison, we implement the MLP with
1 hidden layer of 500 ReLU units, and the CNN with two convolution layers of
128 feature maps each followed by 2 fully-connected layers of 500 and 10 ReLU
units. The input frames to the MLP and CNN are concatenated over time into a
spectrogram image. Since the number of frames for each sound clip varies from 20
to 100 and cannot be processed directly by the MLP or CNN, we bilinearly rescale
these spectrogram images into a consistent dimension of 20 × 64.
We implement both the RNN and LSTM with two hidden layers containing 100
hidden units each, and a dropout layer with a probability of 0.5 is applied after the
first hidden layer to prevent overfitting. The input to the RNN and LSTM are the 20dimensional filter bank output vectors. The weights for all networks are initialized
with orthogonal conditions as suggested in [37]. The deep learning networks are
trained with the cross-entropy criterion and optimized using the Adam [38] optimizer.
The learning rate is decayed to 99% of the original value after every epoch, and all
networks are trained for 100 epochs, except for the CNN (50 epochs), by when
convergence is observed. Simulations are repeated 10 times for each model, with
random weight initialization.
3.1.4 Noise Robustness Evaluation
Environmental Noise We generate noise-corrupted sound samples by adding
“Speech Babble” background noise from the NOISEX-92 dataset [39] to the clean
RWCP sound samples. This selected background noise represents a non-stationary
noisy environment with predominantly low-frequency contents, hence making a fair
comparison with the noise robustness tests performed in LSF-SNN [26] and LTFSNN models [27]. For each training or testing sound sample, a random noise segment
of the same duration is selected from the noise file and added at 4 different SNR levels
of 20, 10, 0 and −5 dB separately, giving a total of 1000 training and 1000 testing
samples. The SNR ratio is calculated based on the energy level of each sound sample
and the corresponding noise segment in our experiments. Training is performed over
the whole training set, while the testing set is evaluated separately at different SNR
levels.
We perform multi-condition training on all the MLP, CNN, RNN, LSTM and
SOM-SNN models. Additionally, we also conduct experiments whereby the models
are trained with clean sound samples but tested with noise-corrupted samples (the
mismatched condition).
425
3.1.3 Traditional Artificial Neural Networks
To facilitate comparison with other traditional ANN models trained on the RWCP
dataset, we implement four common neural network architectures, namely the MultiLayer Perceptron (MLP) [33], the Convolutional Neural Network (CNN) [34], the
Recurrent Neural Network (RNN) [35] and the Long Short-Term Memory (LSTM)
[36] using the Pytorch library. For a fair comparison, we implement the MLP with
1 hidden layer of 500 ReLU units, and the CNN with two convolution layers of
128 feature maps each followed by 2 fully-connected layers of 500 and 10 ReLU
units. The input frames to the MLP and CNN are concatenated over time into a
spectrogram image. Since the number of frames for each sound clip varies from 20
to 100 and cannot be processed directly by the MLP or CNN, we bilinearly rescale
these spectrogram images into a consistent dimension of 20 × 64.
We implement both the RNN and LSTM with two hidden layers containing 100
hidden units each, and a dropout layer with a probability of 0.5 is applied after the
first hidden layer to prevent overfitting. The input to the RNN and LSTM are the 20dimensional filter bank output vectors. The weights for all networks are initialized
with orthogonal conditions as suggested in [37]. The deep learning networks are
trained with the cross-entropy criterion and optimized using the Adam [38] optimizer.
The learning rate is decayed to 99% of the original value after every epoch, and all
networks are trained for 100 epochs, except for the CNN (50 epochs), by when
convergence is observed. Simulations are repeated 10 times for each model, with
random weight initialization.
3.1.4 Noise Robustness Evaluation
Environmental Noise We generate noise-corrupted sound samples by adding
“Speech Babble” background noise from the NOISEX-92 dataset [39] to the clean
RWCP sound samples. This selected background noise represents a non-stationary
noisy environment with predominantly low-frequency contents, hence making a fair
comparison with the noise robustness tests performed in LSF-SNN [26] and LTFSNN models [27]. For each training or testing sound sample, a random noise segment
of the same duration is selected from the noise file and added at 4 different SNR levels
of 20, 10, 0 and −5 dB separately, giving a total of 1000 training and 1000 testing
samples. The SNR ratio is calculated based on the energy level of each sound sample
and the corresponding noise segment in our experiments. Training is performed over
the whole training set, while the testing set is evaluated separately at different SNR
levels.
We perform multi-condition training on all the MLP, CNN, RNN, LSTM and
SOM-SNN models. Additionally, we also conduct experiments whereby the models
are trained with clean sound samples but tested with noise-corrupted samples (the
mismatched condition).
