424
J. Wu et al.
from each class, of which 20 are used for training and the remaining 20 for testing,
giving a total of 200 training and 200 testing samples.
The TIDIGITS [32] dataset consists of reading digit strings of varying lengths, and
speech signals are sampled at 20 kHz. The TIDIGITS dataset is a publicly available
dataset from the Linguistic Data Consortium, which is one of the most commonly
available speech datasets used for benchmarking speech recognition algorithms. This
dataset consists of spoken digit utterances from 111 male and 114 female speakers.
We used all of the 12,373 continuous spoken digit utterances for the SOM training and
the rest of the 4950 isolated spoken digit utterances for the SNN training and testing.
Each speaker contributes two isolated spoken digit utterances for all 11 classes (i.e.,
‘zeros’ to ‘nine’ and ‘oh’). We split the isolated spoken digit utterances randomly
with 3950 utterances for training and the remaining 1000 utterances for testing.
3.1.2 SOM-SNN Framework
The SOM-SNN framework, as shown in Fig. 1, consists of three processing stages
organized in a pipeline. These stages are trained separately and then evaluated in a
single, continuous process. For the auditory front-end, we segment the continuous
sound samples into frames of 100 ms length with 50 ms overlap between neighboring
frames for the RWCP dataset. In contrast, we use a frame length of 25 ms with 10
ms overlap for the TIDIGITS dataset. These values are determined empirically to
sufficiently discriminate the signals without excessive computational load. We utilize
20 mel-scaled filters for the spectral analysis, ranging from 200 to 8000 Hz and 200 to
10,000 Hz respectively for the RWCP and TIDIGITS datasets. The number of filters
is again empirically determined, such that more filters do not improve classification
accuracy.
For feature representation learning in the SOM, we utilize the SOM available in the
MATLAB Neural Network Toolbox. The Euclidean distance is used to determine
the BMUs, which are subsequently converted into spatiotemporal spike patterns.
The output spikes from the SOM are generated per sound frame, with an interval as
determined by the frame shift (i.e., 50 ms for RWCP dataset and 15 ms for TIDIGITS
dataset).
We initialize the SNN by setting the threshold V thr , the hard margin and learning
rate λ to 1.0, 0.5 and 0.005, respectively. The time constants of the SNN have
determined empirically such that the PSP duration is optimal for the particular dataset,
and we set τ m to 750, 225 ms and τ s to 187.5, 56.25 ms for the RWCP and TIDIGITS
datasets, respectively. We train all the SNNs for 10 epochs by when convergence
is observed. The initial weights for the neurons in the SNN classifier are drawn
randomly from the Gaussian distribution with a mean of 0 and standard deviation of
10
−3 . Parameters used in all our experiments are as above unless otherwise stated.
Précédent

- 424/439

Suivant