bioRxiv · 10.64898/2026.02.06.704508
NeuroVLM: A generative vision-language framework for human neuroimaging
Abstract
Neuroimaging research has produced tens-of-thousands of articles that pair natural language and activation coordinate tables. Recent advances in vision-language models (VLMs) have provided methods to model text and images simultaneously. In this work, we present NeuroVLM, a model architecture for learning from 30,826 human neuroimage-text pairs. The architecture supports contrastive and generative objectives. The contrastive model ranks similarity between neuroimages and text. The generative models include text-to-neuroimage and neuroimage-to-text. These models are evaluated on network images from a variety of atlases, statistical maps from diverse publications, and images created from coordinate tables. These models are capable of generating atlases or maps given a text corpus, generating text interpretations of neuroimages, labeling networks, finding publications most related to a neuroimage query, or finding neuroimages most related to a text query.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hammonds, R. P., Aguirre-Chavez, J., Omoma-Edosa, B., Voytek, B. P.. 2026-02-09. NeuroVLM: A generative vision-language framework for human neuroimaging. https://doi.org/10.64898/2026.02.06.704508
Cite the original work for its findings. Save a collection to share your selection of sources.