An Automatic Sound ClassifiCation Framework …
421
K (t − t j ) = K 0
exp(−
t − t j
τ m
) − exp(−
t − t j
τ s
)
θ(t − t j )
(7)
where K 0 is a normalization factor that ensures the maximum value of the kernel
K (t − t j ) is 1. τ m and τ s correspond to the membrane and synaptic time constants,
which jointly determine the shape of the kernel function. In addition, θ(t − t j ) represents the Heaviside function to ensure that only pre-synaptic spikes emitted before
time t are considered.
θ(x) =
1, i f x ≥ 0
0, other wise
(8)
At time t, the membrane potential of the post-synaptic neuron i is determined by the
weighted sum of all PSPs triggered by incoming spikes before time t:
V i (t) =
j
w ji
t j K (t − t j ) + V rest ∀t ∈ [0, T ]
(9)
where w ji is the synaptic weight between the pre-synaptic neuron j and post-synaptic
neuron i, and T is the duration of the simulation. Whenever the membrane potential
V i (t) of the post-synaptic neuron i reaches the firing threshold, it emits a spike. For
the single-spike based classifier used in this work, the membrane potential of the
post-synaptic neuron then smoothly relaxes back to V rest after spiking by shunting
all subsequent input spikes (i.e., input spikes arriving after the post-synaptic spike
have no effect on the membrane potential of the post-synaptic neuron). Since these
input spikes would not contribute to any learning in the single-spike based classifier,
the unnecessary post-spike computations can be safely ignored.
2.3.2 Maximum-Margin Tempotron Learning Rule
For the classification of spatiotemporal patterns as illustrated by the SNN in Fig. 1,
we use a modified version of the biologically plausible Tempotron [25] learning rule
to train the classifier, which has been successfully used in several ASC tasks [19,
26–28]. The original Tempotron rule is designed for a binary classification task, such
that a neuron emits a spike when it observes a spike pattern from its desired class,
and remains quiescent otherwise. For a multi-class classification task, we adopt the
one-against-all strategy to train one output neuron to respond to each class.
During training, for neuron i that represents the ith class, we treat all training
samples with class label i as positive samples, and all others as negative. During
testing, we monitor the membrane potential of all output neurons and classify the
test sample as follows: (1) If no output neuron fires over the sample duration, we
select the output neuron with the highest membrane potential as the correct class. (2)
If only a single output neuron fires, the class label corresponding to this neuron is
selected. (3) Otherwise, if two or more neurons fire, we label the test sample with the
Précédent

- 421/439

Suivant