418
J. Wu et al.
2.1 Auditory Front-End
Human auditory front-end consists of the outer, middle and inner ear. In the outer
ear, sound waves travel through air and arrive at the pinna, which also embeds the
location information of the sound source. From the pinna, the sound signals are then
transmitted via the ear canal, which functions as a resonator, to the middle ear. In the
middle ear, vibrations (induced by the sound signals) are converted into mechanical
movements of the ossicles (i.e., malleus, incus, and stapes) through the tympanic
membrane. The tensor tympani and stapedius muscles, which are connected to the
ossicles, act as an automatic gain controller to moderate mechanical movements
under the high-intensity scenario. At the end of the middle ear, the ossicles join with
the cochlea via the oval window, where mechanical movements of the ossicles are
transformed into fluid pressure oscillations which move along the basilar membrane
in the cochlea [20].
The cochlea is a wonderful anatomical work of art. It functions as a spectrum
analyzer which displaces the basilar membrane at specific locations that correspond
to different frequency components in the sound wave. Finally, displacements of
the basilar membrane activate inner hair cells via nearby mechanically gated ion
channels, converting mechanical displacements into electrical impulse trains. The
spike trains generated at the hair cells are transmitted to the cochlear nuclei through
dedicated auditory nerves. Functionally, the cochlear nuclei act as filter banks, which
also normalize activities of saturated auditory nerve fibers over different frequency
bands. Most of the auditory nerves terminate at the cochlear nuclei where sound
information is still identifiable. Beyond the cochlear nuclei, in the auditory cortex,
it remains unclear how information is being represented and processed [21].
The understanding of the human auditory front-end has a significant impact on
machine hearing research and inspires many biologically plausible feature representations of acoustic signals, such as the MFCC and GTCC. Here, we adopt the
MFCC representation. As shown in Fig. 1, we pre-processed the sound signals by
first applying pre-emphasis to amplify high-frequency contents, then segmenting the
continuous sound signals into overlapped frames of suitable length so as to better
capture the temporal variations of the sound signal, and finally applying the Hamming window on these frames to reduce the effect of spectral leakage. To extract the
spectral contents in the acoustic stimuli, we perform Short-Time Fourier-Transform
(STFT) on the sound frames and compute the power spectrum. After that, we apply
20 logarithmic mel-scaled filters on the resulting power spectrum, generating a compressed feature representation for each sound frame. The mel-scaled filter banks
emulate the human perception of sound that is more discriminative towards the low
frequency as compared to the high frequency components.
Précédent

- 418/439

Suivant