Search bioRxiv⌕ Search

Biology subjects

Khalilian-Gourtani, A.

Publications and source records attributed to Khalilian-Gourtani, A..

3 recordsLinked to original sources

A Neural Speech Decoding Framework Leveraging Deep Learning and Speech Synthesis

Decoding human speech from neural signals is essential for brain-computer interface (BCI) technologies restoring speech function in populations with neurological deficits. However, it remains a highly challenging task, compounded by the scarce availability of neural signals with corresponding speech, data complexity, and high dimensionality, and the limited publicly available source code. Here, we present a novel deep learning-based neural speech decoding framework that includes an ECoG Decoder that translates electrocorticographic (ECoG) signals from the cortex into interpretable speech parameters and a novel differentiable Speech Synthesizer that maps speech parameters to spectrograms. We develop a companion audio-to-audio auto-encoder consisting of a Speech Encoder and the same Speech Synthesizer to generate reference speech parameters to facilitate the ECoG Decoder training. This framework generates natural-sounding speech and is highly reproducible across a cohort of 48 participants. Among three neural network architectures for the ECoG Decoder, the 3D ResNet model has the best decoding performance (PCC=0.804) in predicting the original speech spectrogram, closely followed by the SWIN model (PCC=0.796). Our experimental results show that our models can decode speech with high correlation even when limited to only causal operations, which is necessary for adoption by real-time neural prostheses. We successfully decode speech in participants with either left or right hemisphere coverage, which could lead to speech prostheses in patients with speech deficits resulting from left hemisphere damage. Further, we use an occlusion analysis to identify cortical regions contributing to speech decoding across our models. Finally, we provide open-source code for our two-stage training pipeline along with associated preprocessing and visualization tools to enable reproducible research and drive research across the speech science and prostheses communities.

neuroscience↗

A Corollary Discharge Circuit in Human Speech

When we vocalize, our brain distinguishes self-generated sounds from external ones. A corollary discharge signal supports this function in animals, however, in humans its exact origin and temporal dynamics remain unknown. We report Electrocorticographic (ECoG) recordings in neurosurgical patients and a novel connectivity approach based on Granger-causality that reveals major neural communications. We find a reproducible source for corollary discharge across multiple speech production paradigms localized to ventral speech motor cortex before speech articulation. The uncovered discharge predicts the degree of auditory cortex suppression during speech, its well-documented consequence. These results reveal the human corollary discharge source and timing with far-reaching implication for speech motor-control as well as auditory hallucinations in human psychosis. Significance statementHow do organisms dissociate self-generated sounds from external ones? A fundamental brain circuit across animals addresses this question by transmitting a blueprint of the motor signal to sensory cortices, referred to as a corollary discharge. However, in humans and non-human primates auditory system, the evidence supporting this circuit has been limited to its direct consequence, auditory suppression. Furthermore, an impaired corollary discharge circuit in humans can lead to auditory hallucinations. While hypothesized to originate in the frontal cortex, direct evidence localizing the source and timing of an auditory corollary discharge is lacking in humans. Leveraging rare human neurosurgical recordings combined with connectivity techniques, we elucidate the exact source and dynamics of the corollary discharge signal in human speech. One-sentence summaryWe reveal the source and timing of a corollary discharge from speech motor cortex onto auditory cortex in human speech.

neuroscience↗

Distributed Feedforward and Feedback Processing across Perisylvian Cortex Supports Human Speech

Speech production is a complex human function requiring continuous feedforward commands together with reafferent feedback processing. These processes are carried out by distinct frontal and posterior cortical networks, but the degree and timing of their recruitment and dynamics remain unknown. We present a novel deep learning architecture that translates neural signals recorded directly from cortex to an interpretable representational space that can reconstruct speech. We leverage state-of-the-art learnt decoding networks to disentangle feedforward vs. feedback processing. Unlike prevailing models, we find a mixed cortical architecture in which frontal and temporal networks each process both feedforward and feedback information in tandem. We elucidate the timing of feedforward and feedback related processing by quantifying the derived receptive fields. Our approach provides evidence for a surprisingly mixed cortical architecture of speech circuitry together with decoding advances that have important implications for neural prosthetics.

neuroscience↗