Search bioRxiv⌕ Search

Biology subjects

Buss, E.

Publications and source records attributed to Buss, E..

3 recordsLinked to original sources

Talker head-orientation and extended high-frequency benefits for speech recognition as a function of masker head angle

Several types of cues contribute to speech recognition in multi-talker environments. In this study, we investigated how talker head-orientation related (THOR) cues and extended high- frequency (EHF; >8kHz) cues affect speech-in-speech recognition for both female and male speech. We examined the THOR benefit associated with a non-facing masker talker head orientation (relative to a facing orientation) as a function of masker talker facing angle. The target talker always faced the listener, whereas co-located maskers were tested with eight different masker head angles, ranging from 0{degrees} (facing the listener) to facing 180{degrees} away. Two filtering conditions were tested: full- band and low-pass filtered at 8 kHz. A THOR benefit was observed at masker head angles greater than 45{degrees}, increasing from 2 dB to 8 dB between angles of 67.5{degrees} and 180{degrees}. This benefit was reduced for low-pass filtered speech. Access to EHF cues improved performance, but only for masker head angles >22.5{degrees}. There was no significant relationship between 16-kHz pure-tone thresholds and performance for young, normal-hearing listeners with good EHF hearing. These findings indicate that listeners benefit from non-facing masker talker head orientations >45{degrees} when the target talker is facing the listener, with greater benefit for larger head angles.

neuroscience↗

FASTIMAGES: Validating replay detection methods in human Neuroimaging using a combined MEG and fMRI dataset

Studies in rodents and humans using invasive electrophysiology have established that neural replay is a ubiquitous phenomenon in the brain that is associated with a wide range of cognitive functions, including memory, planning and decision making. Yet, invasively recording in humans remains difficult, and hence knowledge about replay in humans remains scarce. Hence, to comprehensively understand replay in humans, we need reliable approaches that can detect it non-invasively. Several main non-invasive approaches have been proposed, but we lack a full comparative validation against known ground truth signals. In this study, we present FASTIMAGES, a benchmark dataset from seventy participants with parallel fMRI (n = 40, previously published) and MEG (n=30) recordings containing known neural sequences evoked by fast visual stimulation as well as functional localizer trials. The neural sequences were elicited by five different visual stimuli shown in sequences at speeds of 132, 164, 228 and 612 milliseconds onset-to-onset intervals. Using this dataset, we investigate two existing statistical methods for sequence detection, namely Temporally Delayed Linear Modelling (TDLM, developed for MEG by Liu et al., 2021) and Slope Order Dynamic Analysis (SODA, developed for fMRI by Wittkuhn & Schuck, 2021). We examine the underlying assumptions of each method, analyse their resulting strengths and weaknesses in application to MEG and fMRI. We demonstrate that both approaches excel in their native modality (TDLM for MEG and SODA for fMRI), with comparable effect sizes given idealized conditions in this benchmark. Cross-modality transfer remains challenging. Finally, the FASTIMAGES dataset provides data with known and clearly expressed sequences and can be used to benchmark and validate future sequence detection methods under idealized conditions.

neuroscience↗

Gender and speech material effects on the long-term average speech spectrum, including at extended high frequencies

Gender and language effects on the long-term average speech spectrum (LTASS) have been reported, but typically using recordings that were bandlimited and/or failed to accurately capture extended high frequencies (EHFs). Accurate characterization of the full-band LTASS is warranted given recent data on the contribution of EHFs to speech perception. The present study characterized the LTASS for high-fidelity, anechoic recordings of males and females producing Bamford-Kowal-Bench (BKB) sentences, digits, and unscripted narratives. Gender had an effect on spectral levels at both ends of the spectrum: males had higher levels than females below approximately 160 Hz, owing to lower fundamental frequencies; females had [~]4 dB higher levels at EHFs, but this effect was dependent on speech material. Gender differences were also observed at [~]300 Hz, and between 800-1000 Hz, as previously reported. Despite differences in phonetic content, there were only small, gender-dependent differences in EHF levels across speech materials. EHF levels were highly correlated across materials, indicating relative consistency within talkers. Our findings suggest that LTASS levels at EHFs are influenced primarily by talker and gender, highlighting the need for future research to assess whether EHF cues are more audible for female speech than for male speech.

biophysics↗