Search bioRxiv⌕ Search

Biology subjects

Riegel, J.

Publications and source records attributed to Riegel, J..

6 recordsLinked to original sources

Deep Learning Reveals Cross-Modal Neural Representations of Auditory and Visual Mental Imagery in MEG

Mental imagery provides a unique window into the brains ability to internally simulate sensory experiences, offering valuable insights for both cognitive neuroscience and brain-computer interface (BCI) research. This study examined the neural representations of imagined auditory and visual stimuli using magnetoen-cephalography (MEG) and assessed the ability of machine learning models to decode these mental processes. MEG data were recorded from 18 right-handed participants during auditory and visual imagery tasks and source-reconstructed within modality-specific cortical regions of interest. We compared a convolutional neural network (CNN) and a linear logistic regression model within a subject-specific classification frame-work. Both approaches achieved above-chance decoding accuracies, with the CNN outperforming the linear model in the auditory task, whereas the linear model showed slightly higher accuracy for visual imagery. Notably, the CNN achieved significant decoding performance even when trained on non-task-relevant cortical regions, indicating that imagined stimuli are represented in distributed and partially overlapping neural networks across modalities. This cross-modal decoding capability highlights the potential of deep learning models to capture complex, multimodal neural patterns and suggests that future brain-computer interfaces could benefit from integrating auditory and visual information. A secondary, behavioral analysis revealed correlation of memory capacity and individual learning preferences with decoding performances, suggesting that individual cognitive differences may further shape the quality of neural representations. Together, these findings advance our understanding of cross-modal mental imagery and point toward more flexible and personalized approaches in BCI design. New and NoteworthyBy comparing linear and deep classifiers, this work shows that convolutional networks capture rich, cross-modal neural representations of auditory and visual mental imagery in MEG. Significant decoding from non-task-relevant regions indicates distributed cortical engagement, highlighting deep learnings potential for robust, modality-independent brain-computer interfaces.

neuroscience↗

The cortical contribution to the speech-FFR is not modulated by visual information

Seeing a speakers face can significantly aid understanding, particularly in challenging acoustic environments. An early neural response implicated in audiovisual speech processing is the frequency-following response (speech-FFR), which occurs at the fundamental frequency of the speech signal. This response arises from both subcortical areas and the auditory cortex. Previous studies have shown that subcortical responses are reduced when bimodal stimulation includes visual input from the talkers face. Here, we examined the cortical contribution to the speech-FFR and its potential modulation by visual information. We recorded MEG responses to four types of audiovisual signals: a still image, an artificially generated avatar, a degraded video, and a natural video. The audio stimuli were presented in a substantial level of background noise to make behavioral audiovisual effects stand out. Speech-in-noise comprehension increased significantly from the audio-only condition to the avatar and the degraded video, and further to the natural video. Moreover, we found that all types of audiovisual stimuli yielded robust speech-FFRs in the auditory cortex at an early latency of around 30 ms. However, the magnitude of this neural response was neither enhanced nor attenuated by the videos, nor could the cortical contribution of the speech-FFR explain a significant portion of the variance in the behavioral comprehension scores. Our results suggest that visual modulation of the speech-FFR in the auditory cortex is, if existent, too small to be measurable in scenarios where speech occurs in considerable background noise.

neuroscience↗

Talking avatars can differentially modulate cortical speech tracking in the high and in the low delta band

In noisy listening environments, visual cues from a speakers face can significantly boost speech compre-hension. The underlying audiovisual integration in the brain involves neural tracking of audiovisual speech features. Moreover, lip reading in silence is associated with tracking of the speech envelope in the low-delta frequency band (0.5 - 1 Hz). Recently, digital avatars have emerged that can support speech comprehen-sion. Yet, it remains unclear how the human brain integrates such artificial visual signals with natural speech. Here, we employed magnetoencephalography (MEG) to measure the neural response to a natural video, an avatar generated by deep neural networks, and a degraded video serving as a control. We demonstrate that the avatar can enhance speech-in-noise comprehension to a similar degree as the degraded video, although less than the natural video. We further identify a late response at 600 ms in the neural tracking of the audi-tory cortex in the high delta band (1 - 4 Hz) that predicts audiovisual speech comprehension. In contrast, we found that neural tracking in the low delta band is related to silent lip-reading performance. Importantly, the tracking in the low delta band evoked by the avatars is much weaker and occurs earlier than that elicited by the other audiovisual stimuli. Neural tracking in the theta band (4 - 8 Hz) is not involved in audiovisual integration. Our results show that the low delta band and the high delta band play clearly distinct roles in visual-only and audiovisual speech processing, and suggest potential avenues for further boosting the abilities of avatars to support speech comprehension. Significance StatementUnderstanding a conversational partner is essential for everyday communication. Yet, many people -- due to aging or other factors -- struggle to follow speech in noisy environments. Seeing the speakers face can greatly enhance speech comprehension, but visual cues are often unavailable, such as during public announcements or telephone conversations. Digital avatars offer a promising alternative, but how the brain integrates audiovisual information from such artificial sources remains unclear. Using magnetoencephalog-raphy (MEG), we investigated how the brain processes and integrates speech when visual information is provided by either natural or artificial (avatar-based) signals. Our findings reveal both shared and distinct neural mechanisms of audiovisual integration, providing critical insight into how visual input can support speech understanding in challenging listening conditions.

neuroscience↗

Attention Decoding at the Cocktail Party: Preserved in Hearing Aid Users, Reduced in Cochlear Implant Users

AbstractUsers of hearing aids (HAs) and cochlear implants (CIs) experience significant difficulty understanding a target speaker in multi-talker environments or when other background noise is present. Segregation of a particular voice from background noise occurs partly through enhanced cortical tracking of amplitude fluctuations in the target signal. Measuring a persons cortical tracking allows decoding their focus of attention and may be used for neurofeedback in hearing devices, potentially aiding their users with speech-in-noise comprehension. Most studies on cortical speech tracking have employed typical hearing (TH) individuals, whereas studies in people with hearing impairment whose cortical tracking may differ are still scarce. The objective of this study was to compare cortical speech tracking of HA (n=29) and CI users (n=24) to that of age-matched TH individuals (n=29). We recorded EEG data while the participants attended one of two competing talkers (one with a female and one with a male voice), in a free-field acoustic environment. Importantly, HA users as well as CI users used their personal, clinically-fitted devices. Cortical speech tracking was assessed through linear backward and forward models that related the EEG data to the speech envelope. For the CI users, electrical artifacts stemming from the implant were addressed through a bespoke method for artifact rejection. We found that the HA group exhibited cortical tracking and attentional modulation that were largely comparable to those of the TH group. CI users also showed successful cortical tracking. However, they displayed a profound deficit in attentional modulation, seen in the significantly poorer neural segregation of the attended vs. the ignored speech streams. These results shed light on a neurobiological mechanism for speech-in-noise comprehension and have implications for neurofeedback in hearing devices.

neuroscience↗

Assessing the Impact of Selective Attention on the Cortical Tracking of the Speech Envelope in the Delta and Theta Frequency Bands and How Musical Training Does (Not) Affect It

Oral communication regularly takes place amidst background noise, requiring the ability to selectively attend to a target speech stream. Musical training has been shown to be beneficial for this task. Regarding the underlying neural mechanisms, recent studies showed that the speech envelope is tracked by neural activity in the auditory cortex, which plays a role in the neural processing of speech, including speech in noise. The neural tracking occurs predominantly in two frequency bands, the delta and the theta band. However, much regarding the specifics of these neural responses, as well as their modulation through musical training, still remain unclear. Here, we investigated the delta- and theta-band cortical tracking of the speech envelope of attended and ignored speech using magnetoencephalography (MEG) recordings. We thereby assessed both musicians and non-musicians to explore potential differences between these groups. The cortical speech tracking was quantified through source-reconstructing the MEG data and subsequently relating the speech envelope in a certain frequency band to the MEG data using linear models. We thereby found the theta-band tracking to be dominated by early responses with comparable magnitudes for attended and ignored speech, whereas the delta band tracking exhibited both earlier and later responses that were modulated by selective attention. Almost no significant differences emerged in the neural responses between musicians and non-musicians. Our findings show that only the speech tracking in the delta but not in the theta band contributes to selective attention, but that this mechanism is essentially unaffected by musical training.

neuroscience↗

No evidence of musical training influencing the cortical contribution to the speech-FFR and its modulation through selective attention

Musicians can have better abilities to understand speech in adverse conditions such as background noise than non-musicians. However, the neural mechanisms behind such enhanced behavioral performances remain largely unclear. Studies have found that the subcortical frequency-following response to the fundamental frequency of speech and its higher harmonics (speech-FFR) may be involved since it is larger in people with musical training than in those without. Recent research has shown that the speech-FFR consists of a cortical contribution in addition to the subcortical sources. Both the subcortical and the cortical contribution are modulated by selective attention to one of two competing speakers. However, it is unknown whether the strength of the cortical contribution to the speech-FFR, or its attention modulation, is influenced by musical training. Here we investigate these issues through magnetoencephalographic (MEG) recordings of 52 subjects (18 musicians, 25 non-musicians, and 9 neutral participants) listening to two competing male speakers while selectively attending one of them. The speech-in-noise comprehension abilities of the participants were not assessed. We find that musicians and non-musicians display comparable cortical speech-FFRs and additionally exhibit similar subject-to-subject variability in the response. Furthermore, we also do not observe a difference in the modulation of the neural response through selective attention between musicians and non-musicians. Moreover, when assessing whether the cortical speech-FFRs are influenced by particular aspects of musical training, no significant effects emerged. Taken together, we did not find any effect of musical training on the cortical speech-FFR. Significance statementIn previous research musicians have been found to exhibit larger subcortical responses to the pitch of a speaker than non-musicians. These larger responses may reflect enhanced pitch processing due to musical training and may explain why musicians tend to understand speech better in noisy environments than people without musical training. However, higher-level cortical responses to the pitch of a voice exist as well and are influenced by attention. We show here that, unlike the subcortical responses, the cortical activities do not differ between musicians and non-musicians. The attentional effects are not influenced by musical training. Our results suggest that, unlike the subcortical response, the cortical response to pitch is not shaped by musical training.

neuroscience↗