Search bioRxiv⌕ Search

Biology subjects

Kolossa, D.

Publications and source records attributed to Kolossa, D..

2 recordsLinked to original sources

The cortical contribution to the speech-FFR is not modulated by visual information

Seeing a speakers face can significantly aid understanding, particularly in challenging acoustic environments. An early neural response implicated in audiovisual speech processing is the frequency-following response (speech-FFR), which occurs at the fundamental frequency of the speech signal. This response arises from both subcortical areas and the auditory cortex. Previous studies have shown that subcortical responses are reduced when bimodal stimulation includes visual input from the talkers face. Here, we examined the cortical contribution to the speech-FFR and its potential modulation by visual information. We recorded MEG responses to four types of audiovisual signals: a still image, an artificially generated avatar, a degraded video, and a natural video. The audio stimuli were presented in a substantial level of background noise to make behavioral audiovisual effects stand out. Speech-in-noise comprehension increased significantly from the audio-only condition to the avatar and the degraded video, and further to the natural video. Moreover, we found that all types of audiovisual stimuli yielded robust speech-FFRs in the auditory cortex at an early latency of around 30 ms. However, the magnitude of this neural response was neither enhanced nor attenuated by the videos, nor could the cortical contribution of the speech-FFR explain a significant portion of the variance in the behavioral comprehension scores. Our results suggest that visual modulation of the speech-FFR in the auditory cortex is, if existent, too small to be measurable in scenarios where speech occurs in considerable background noise.

neuroscience↗

Talking avatars can differentially modulate cortical speech tracking in the high and in the low delta band

In noisy listening environments, visual cues from a speakers face can significantly boost speech compre-hension. The underlying audiovisual integration in the brain involves neural tracking of audiovisual speech features. Moreover, lip reading in silence is associated with tracking of the speech envelope in the low-delta frequency band (0.5 - 1 Hz). Recently, digital avatars have emerged that can support speech comprehen-sion. Yet, it remains unclear how the human brain integrates such artificial visual signals with natural speech. Here, we employed magnetoencephalography (MEG) to measure the neural response to a natural video, an avatar generated by deep neural networks, and a degraded video serving as a control. We demonstrate that the avatar can enhance speech-in-noise comprehension to a similar degree as the degraded video, although less than the natural video. We further identify a late response at 600 ms in the neural tracking of the audi-tory cortex in the high delta band (1 - 4 Hz) that predicts audiovisual speech comprehension. In contrast, we found that neural tracking in the low delta band is related to silent lip-reading performance. Importantly, the tracking in the low delta band evoked by the avatars is much weaker and occurs earlier than that elicited by the other audiovisual stimuli. Neural tracking in the theta band (4 - 8 Hz) is not involved in audiovisual integration. Our results show that the low delta band and the high delta band play clearly distinct roles in visual-only and audiovisual speech processing, and suggest potential avenues for further boosting the abilities of avatars to support speech comprehension. Significance StatementUnderstanding a conversational partner is essential for everyday communication. Yet, many people -- due to aging or other factors -- struggle to follow speech in noisy environments. Seeing the speakers face can greatly enhance speech comprehension, but visual cues are often unavailable, such as during public announcements or telephone conversations. Digital avatars offer a promising alternative, but how the brain integrates audiovisual information from such artificial sources remains unclear. Using magnetoencephalog-raphy (MEG), we investigated how the brain processes and integrates speech when visual information is provided by either natural or artificial (avatar-based) signals. Our findings reveal both shared and distinct neural mechanisms of audiovisual integration, providing critical insight into how visual input can support speech understanding in challenging listening conditions.

neuroscience↗