An Automatic Sound ClassifiCation Framework …
419
2.2 Feature Representation Using SOM
Feature representation is critical in all ASC systems; state-of-the-art ASC systems
input low-level MFCC or GTCC features into the GMM-HMM or deep learning
models so as to extract higher-level representations. In our initial experiments, we
observe that existing SNN temporal learning rules cannot discriminate latency [22]
or population [23] encoded mel-scaled filter bank outputs effectively. Therefore, we
propose to use the biologically inspired SOM to form a mid-level feature representation of the sound frames. The neurons in the SOM form distinctive synaptic filters
that organize themselves tonotopically and compete to represent the filter bank output
vectors. Such tonotopically organized feature maps have been found in the human
auditory cortex in many physiological experiments [24].
As shown in Fig. 1, all neurons in the SOM are fully connected to the filter bank
and receive mel-scaled filter outputs (real-valued vectors). The SOM learns acoustic
features in an unsupervised manner, whereby two mechanisms: competition and
cooperation, guide the formation of a tonotopically organized neural map. During
training, the neurons in the SOM compete with each other to best represent the input
frame. The best-matching unit (BMU), with its synaptic weight vector closest to the
input vector in the feature space, will update its weight vector to become closer to
the input vector. Additionally, the neurons surrounding the BMU will cooperate with
it by updating their weight vectors to move closer to the input vector. The magnitude
of the weight update of neighboring neurons is inversely proportional to its distance
to the BMU, effectively facilitating the formation of neural clusters. Eventually, the
synaptic weight vectors of neurons in the SOM follow the distribution of input feature
vectors and organize tonotopically, such that adjacent neurons in the SOM will have
similar weight vectors.
During the evaluation, as shown in Fig. 1, the SOM (through the BMU neuron)
emits a single spike at each sound frame sampling interval. The sparsely activated
BMUs encourage pattern separation and enhance power efficiency. The spikes triggered over the duration of a sound event form a spatiotemporal spike pattern, which
is then classified by the SNN into one of the sound classes. The mechanisms of
SOM training and testing are provided in Algorithm 1 (for more details see [18]).
This seminal work [18] trained the SOM for a phoneme recognition task, which
then used a set of hand-crafted rules to link sound clusters of the SOM to actual
phoneme classes. Here, we use an SNN-based classifier to automatically categorize
the spatiotemporal spike patterns into different sound events.
419
2.2 Feature Representation Using SOM
Feature representation is critical in all ASC systems; state-of-the-art ASC systems
input low-level MFCC or GTCC features into the GMM-HMM or deep learning
models so as to extract higher-level representations. In our initial experiments, we
observe that existing SNN temporal learning rules cannot discriminate latency [22]
or population [23] encoded mel-scaled filter bank outputs effectively. Therefore, we
propose to use the biologically inspired SOM to form a mid-level feature representation of the sound frames. The neurons in the SOM form distinctive synaptic filters
that organize themselves tonotopically and compete to represent the filter bank output
vectors. Such tonotopically organized feature maps have been found in the human
auditory cortex in many physiological experiments [24].
As shown in Fig. 1, all neurons in the SOM are fully connected to the filter bank
and receive mel-scaled filter outputs (real-valued vectors). The SOM learns acoustic
features in an unsupervised manner, whereby two mechanisms: competition and
cooperation, guide the formation of a tonotopically organized neural map. During
training, the neurons in the SOM compete with each other to best represent the input
frame. The best-matching unit (BMU), with its synaptic weight vector closest to the
input vector in the feature space, will update its weight vector to become closer to
the input vector. Additionally, the neurons surrounding the BMU will cooperate with
it by updating their weight vectors to move closer to the input vector. The magnitude
of the weight update of neighboring neurons is inversely proportional to its distance
to the BMU, effectively facilitating the formation of neural clusters. Eventually, the
synaptic weight vectors of neurons in the SOM follow the distribution of input feature
vectors and organize tonotopically, such that adjacent neurons in the SOM will have
similar weight vectors.
During the evaluation, as shown in Fig. 1, the SOM (through the BMU neuron)
emits a single spike at each sound frame sampling interval. The sparsely activated
BMUs encourage pattern separation and enhance power efficiency. The spikes triggered over the duration of a sound event form a spatiotemporal spike pattern, which
is then classified by the SNN into one of the sound classes. The mechanisms of
SOM training and testing are provided in Algorithm 1 (for more details see [18]).
This seminal work [18] trained the SOM for a phoneme recognition task, which
then used a set of hand-crafted rules to link sound clusters of the SOM to actual
phoneme classes. Here, we use an SNN-based classifier to automatically categorize
the spatiotemporal spike patterns into different sound events.
