Search bioRxiv⌕ Search

Biology subjects

Fonteneau, E.

Publications and source records attributed to Fonteneau, E..

2 recordsLinked to original sources

On the similarities of representations in artificial and brain neural networks for speech recognition

How the human brain supports speech comprehension is an important question in neuroscience. Studying the neurocomputational mechanisms underlying human language is not only critical to understand and develop treatments for many human conditions that impair language and communication but also to inform artificial systems that aim to automatically process and identify natural speech. In recent years, intelligent machines powered by deep learning have achieved near human level of performance in speech recognition. The fields of artificial intelligence and cognitive neuroscience have finally reached a similar phenotypical level despite of their huge differences in implementation, and so deep learning models can--in principle--serve as candidates for mechanistic models of the human auditory system. Utilizing high-performance automatic speech recognition systems, and advanced noninvasive human neuroimaging technology such as magnetoencephalography and multivariate pattern-information analysis, the current study aimed to relate machine-learned representations of speech to recorded human brain representations of the same speech. In one direction, we found a quasi-hierarchical functional organisation in human auditory cortex qualitatively matched with the hidden layers of deep neural networks trained in an automatic speech recognizer. In the reverse direction, we modified the hidden layer organization of the artificial neural network based on neural activation patterns in human brains. The result was a substantial improvement in word recognition accuracy and learned speech representations. We have demonstrated that artificial and brain neural networks can be mutually informative in the domain of speech recognition. Author summaryThe human capacity to recognize individual words from the sound of speech is a cornerstone of our ability to communicate with one another, yet the processes and representations underlying it remain largely unknown. Software systems for automatic speech-to-text provide a plausible model for how speech recognition can be performed. In this study, we used an automatic speech recogniser model to probe recordings from the brains of participants who listened to speech. We found that the parts of the dynamic, evolving representations inside the machine system were a good fit for representations found in the brain recordings, both showing similar hierarchical organisations. Then, we observed where the machines representations diverged from the brains, and made experimental adjustments to the automatic recognizers design so that its representations might better fit the brains. In so doing, we substantially improved the recognizers ability to accurately identify words.

neuroscience↗

Investigating brain mechanisms underlying natural reading by co-registering eye tracking with combined EEG and MEG

Linking brain and behavior is one of the great challenges in cognitive neuroscience. Ultimately, we want to understand how the brain processes information to guide every-day behavior. However, most neuroscientific studies employ very simplistic experimental paradigms whose ecological validity is doubtful. Reading is a case in point, since most neuroscientific studies to date have used unnatural word-by-word stimulus presentation and have often focused on single word processing. Previous research has therefore actively avoided factors that are important for natural reading, such as rapid self-paced voluntary saccadic eye movements. Recent methodological developments have made it possible to deal with associated problems such as eye movement artefacts and the overlap of brain responses to successive stimuli, using a combination of eye-tracking and neuroimaging. A growing number of electroencephalography (EEG) and functional magnetic resonance imaging (fMRI) are successfully using this methodology. Here, we provide a proof-of-concept that this methodology can be applied to combined EEG and magnetoencephalography (MEG) data. Our participants naturally read 4-word sentences that could end in a plausible or implausible word while eye-tracking, EEG and MEG were being simultaneously recorded. Eye-movement artefacts were removed using independent-component analysis. Fixation-related potentials and fields for sentence-final words were subjected to minimum-norm source estimation. We detected an N400-type brain response in our EEG data starting around 200 ms after fixation of the sentence-final word. The brain sources of this effect, estimated from combined EEG and MEG data, were mostly located in left temporal lobe areas. We discuss the possible use of this method for future neuroscientific research on language and cognition.

neuroscience↗