Search bioRxivSearch

Biology subjects

Mesgarani, N.

Publications and source records attributed to Mesgarani, N..

2 recordsLinked to original sources

Phase resetting in human auditory cortex to visual speech

Natural conversation is multisensory: when we can see the speakers face, visual speech cues influence our perception of what is being said. The neuronal basis of this phenomenon remains unclear, though there is indication that phase modulation of neuronal oscillations--ongoing excitability fluctuations of neuronal populations in the brain--provides a mechanistic contribution. Investigating this question using naturalistic audiovisual speech with intracranial recordings in humans, we show that neuronal populations in auditory cortex track the temporal dynamics of unisensory visual speech using the phase of their slow oscillations and phase-related modulations in high-frequency activity. Auditory cortex thus builds a representation of the speech streams envelope based on visual speech alone, at least in part by resetting the phase of its ongoing oscillations. Phase reset could amplify the representation of the speech stream and organize the information contained in neuronal activity patterns. SIGNIFICANCE STATEMENTWatching the speaker can facilitate our understanding of what is being said. The mechanisms responsible for this influence of visual cues on the processing of speech remain incompletely understood. We studied those mechanisms by recording the human brains electrical activity through electrodes implanted surgically inside the skull. We found that some regions of cerebral cortex that process auditory speech also respond to visual speech even when it is shown as a silent movie without a soundtrack. This response can occur through a reset of the phase of ongoing oscillations, which helps augment the response of auditory cortex to audiovisual speech. Our results contribute to discover the mechanisms by which the brain merges auditory and visual speech into a unitary perception.

neuroscience

Reconstructing intelligible speech from the human auditory cortex

Auditory stimulus reconstruction is a technique that finds the best approximation of the acoustic stimulus from the population of evoked neural activity. Reconstructing speech from the human auditory cortex creates the possibility of a speech neuroprosthetic to establish a direct communication with the brain and has been shown to be possible in both overt and covert conditions. However, the low quality of the reconstructed speech has severely limited the utility of this method for brain-computer interface (BCI) applications. To advance the state-of-the-art in speech neuroprosthesis, we combined the recent advances in deep learning with the latest innovations in speech synthesis technologies to reconstruct closed-set intelligible speech from the human auditory cortex. We investigated the dependence of reconstruction accuracy on linear and nonlinear (deep neural network) regression methods and the acoustic representation that is used as the target of reconstruction, including auditory spectrogram and speech synthesis parameters. In addition, we compared the reconstruction accuracy from low and high neural frequency ranges. Our results show that a deep neural network model that directly estimates the parameters of a speech synthesizer from all neural frequencies achieves the highest subjective and objective scores on a digit recognition task, improving the intelligibility by 65% over the baseline method which used linear regression to reconstruct the auditory spectrogram. These results demonstrate the efficacy of deep learning and speech synthesis algorithms for designing the next generation of speech BCI systems, which not only can restore communications for paralyzed patients but also have the potential to transform human-computer interaction technologies.

neuroscience