Search bioRxiv⌕ Search

Biology subjects

Iimura, Y.

Publications and source records attributed to Iimura, Y..

3 recordsLinked to original sources

Speech Synthesis from Electrocorticogram During Imagined Speech Using a Transformer-Based Decoder and Pretrained Vocoder

Synthesizing speech from Electrocorticogram (ECoG) signals recorded during imagined speech remains a challenge due to the absence of synchronized audio signals for training. To address this, we propose a training framework that utilizes audio recorded during overt speech tasks as a surrogate ground truth for imagined speech signals, based on the consistency of the linguistic content. We employed a Transformer-based decoder to generate log-mel spectrograms from imagined speech ECoG, which were then converted into waveform audio using a pre-trained Parallel WaveGAN. In experiments involving ECoG recordings from 13 participants, the synthesized speech achieved dynamic time warping-aligned Pearson correlation coefficients ranging from 0.74 to 0.84 with the proxy targets. These results demonstrate that overt speech audio can serve as an effective training target for reconstructing imagined speech, offering a viable solution for training decoders in the absence of behavioral output.

neuroscience↗

Image retrieval based on closed-loop visual-semantic neural decoding

Neural decoding via the latent space of deep neural network models can infer perceived and imagined images from neural activities, even when the image is novel for the subject and decoder. Brain-computer interfaces (BCIs) using the latent space enable a subject to retrieve intended image from a large dataset on the basis of their neural activities but have not yet been realized. Here, we used neural decoding in a closed-loop condition to retrieve images of the instructed categories from 2.3 million images on the basis of the latent vector inferred from electrocorticographic signals of visual cortices. Using a latent space of contrastive language-image pretraining (CLIP) model, two subjects retrieved images with significant accuracy exceeding 80% for two instructions. In contrast, the image retrieval failed using the latent space of another model, AlexNet. In another task to imagine an image while viewing a different image, the imagery made the inferred latent vector significantly closer to the vector of the imagined category in the CLIP latent space but significantly further away in the AlexNet latent space, although the same electrocorticographic signals from nine subjects were decoded. Humans can retrieve the intended information via a closed-loop BCI with an appropriate latent space.

neuroscience↗

Feasibility of decoding covert speech in ECoG with aTransformer trained on overt speech

Several attempts for speech brain-computer interfacing (BCI) have been made to decode phonemes, sub-words, words, or sentences using invasive measurements, such as the electrocorticogram (ECoG), during auditory speech perception, overt speech, or imagined (covert) speech. Decoding sentences from covert speech is a challenging task. Sixteen epilepsy patients with intracranially implanted electrodes participated in this study, and ECoGs were recorded during overt speech and covert speech of eight Japanese sentences, each consisting of three tokens. In particular, Transformer neural network model was applied to decode text sentences from covert speech, which was trained using ECoGs obtained during overt speech. We first examined the proposed Transformer model using the same task for training and testing, and then evaluated the models performance when trained with overt task for decoding covert speech. The Transformer model trained on covert speech achieved an average token error rate (TER) of 46.6% for decoding covert speech, whereas the model trained on overt speech achieved a TER of 46.3% (p > 0.05; d = 0.07). Therefore, the challenge of collecting training data for covert speech can be addressed using overt speech. The performance of covert speech can improve by employing several overt speeches.

neuroscience↗