bioRxiv · 10.1101/2025.10.22.684047
CellTok: Early-Fusion Multimodal Large Language Model for Single-Cell Transcriptomics via Tokenization
Abstract
Large language models (LLMs) can process diverse forms of information once they are represented as tokens in a shared sequence space. However, single-cell transcriptomes remain a foreign modality to LLMs because they are continuous, high-dimensional molecular profiles rather than discrete linguistic units. Here, we propose CellTok, a tokenized single-cell language modeling approach that converts transcriptomic profiles into compact cellular token sequences and incorporates them into the vocabulary of a pretrained LLM. By representing cells as native tokens, CellTok enables cellular measurements, textual instructions, biological context, and multi-cell populations to be jointly processed within the same autoregressive modeling framework. Across diverse tasks, CellTok enable LLMs to recognize individual cells, interpret homogeneous and heterogeneous cell populations, infer disease-associated cellular states, predict cell-cell communication, model developmental trajectories, and generate cellular states. Moreover, prompt-based experiments show that providing appropriate biological context improves performance, indicating that CellTok can leverage LLM knowledge and contextual reasoning to support cellular data interpretation. These results demonstrate that single-cell transcriptomes can be transformed from a foreign molecular modality into a native language for LLMs, establishing a unified interface for modeling cells, populations, and biological knowledge in a shared token space.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiao, C., Bian, H., Chen, Y., Wei, L., Zhang, X.. 2025-10-24. CellTok: Early-Fusion Multimodal Large Language Model for Single-Cell Transcriptomics via Tokenization. https://doi.org/10.1101/2025.10.22.684047
Cite the original work for its findings. Save a collection to share your selection of sources.