Search bioRxiv⌕ Search

Biology subjects

Clonan, A. C.

Publications and source records attributed to Clonan, A. C..

2 recordsLinked to original sources

Modulation statistics allow robust prediction of speech recognition accuracy across many words, voices, and natural background sounds.

Although humans excel at speech recognition, recognition accuracy can vary widely due to differences in background environments as well as the speakers voice quality, intonation, and pitch. Predicting when speech recognition will succeed or fail, however, remains an ongoing challenge in hearing research. Here we characterize recognition abilities across a wide range of natural conditions using digits spoken by many male and female talkers of multiple ages with 33 unique backgrounds. Across this diverse set of sounds, speech recognition is most strongly influenced by the spectrum and modulation statistics of the noise. Yet, articulatory features of the speech, including fundamental and formant frequencies, show categorically distinct modulatory effects on accuracy across age, gender, and words. We then show that a low-dimensional model of sound, based on computations in the auditory midbrain, accounts for participants single-trial recognition behavior across voices, words and backgrounds. Thus, speech-in-noise perception across extremely diverse natural conditions depends largely on a simple set of spectrotemporal statistics likely encoded by central neural populations.

neuroscience↗

Low-dimensional interference of mid-level sound statistics predicts human speech recognition in natural environmental noise

Recognizing speech in noise, such as in a busy restaurant, is an essential cognitive skill where the task difficulty varies across environments and noise levels. Although there is growing evidence that the auditory system relies on statistical representations for perceiving 1-5 and coding4,6-9 natural sounds, its less clear how statistical cues and neural representations contribute to segregating speech in natural auditory scenes. We demonstrate that human listeners rely on mid-level statistics to segregate and recognize speech in environmental noise. Using natural backgrounds and variants with perturbed spectro-temporal statistics, we show that speech recognition accuracy at a fixed noise level varies extensively across natural backgrounds (0% to 100%). Furthermore, for each background the unique interference created by summary statistics can mask or unmask speech, thus hindering or improving speech recognition. To identify the neural coding strategy and statistical cues that influence accuracy, we developed generalized perceptual regression, a framework that links summary statistics from a neural model to word recognition accuracy. Whereas a peripheral cochlear model accounts for only 60% of perceptual variance, summary statistics from a mid-level auditory midbrain model accurately predicts single trial sensory judgments, accounting for more than 90% of the perceptual variance. Furthermore, perceptual weights from the regression framework identify which statistics and tuned neural filters are influential and how they impact recognition. Thus, perception of speech in natural backgrounds relies on a mid-level auditory representation involving interference of multiple summary statistics that impact recognition beneficially or detrimentally across natural background sounds. Significance StatementRecognizing speech in natural auditory scenes with competing talkers and environmental noise is a critical cognitive skill. Although normal listeners effortlessly perform this task, for instance in a crowded restaurant, it challenges individuals with hearing loss and our most sophisticated machine systems. We tested human participants listening to speech in natural noises with varied statistical characteristics and demonstrate that they rely on a statistical representation of sounds to segregate speech from environmental noise. Using a model of the auditory system, we then demonstrate that a brain inspired statistical representation of natural sounds accurately predicts human perceptual trends across wide range of natural backgrounds and noise levels and reveals key statistical features and neural computations underlying human abilities for this task.

neuroscience↗