bioRxiv · 10.1101/2024.02.16.580578
From Sight to Insight: A Multi-task Approach with the Visual Language Decoding Model
Abstract
Visual neural decoding aims to unlock the mysteries of how the human brain interprets the visual world. While early studies made some progress in decoding visual activity for singular type of information, they failed to concurrently reveal the multi-level interweaving linguistic information in the brain. Here, we developed a novel Visual Language Decoding Model (VLDM) capable of decoding categories, semantic labels, and textual descriptions from visual perceptual activities simultaneously. We selected the large-scale NSD dataset to ensure the efficiency of the decoding model in joint training and evaluation across multiple tasks. For category decoding, we achieved the effective classification of 12 categories with an accuracy of nearly 70%, significantly surpassing the chance level. For label decoding, we attained the precise prediction of 80 specific semantic labels with a 16-fold improvement over the chance level. For text decoding, the scores of the decoded text surpassed the corresponding baseline levels by remarkable margins on six evaluation metrics. This study contributes significantly to extensive applications in multi-layered brain-computer interfaces, potentially leading to more natural and efficient human-computer interaction experiences.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huang, W., Yang, P., Tang, Y., Qin, F., Li, H., Wu, D., Ren, W., Wang, S., Zhao, Y., Wang, J., Liu, H., Li, J., Zhu, Y., Zhou, B., Sun, J., Li, Q., Cheng, K., Yan, H., Chen, H.. 2024-02-20. From Sight to Insight: A Multi-task Approach with the Visual Language Decoding Model. https://doi.org/10.1101/2024.02.16.580578
Cite the original work for its findings. Save a collection to share your selection of sources.