bioRxiv · 10.1101/2024.01.25.577152
scMulan: a multitask generative pre-trained language model for single-cell analysis
Abstract
Cells can be viewed as complex stories written by coordinated expression of genes. The success of AI large language models (LLMs) in mastering the human language inspired us to develop a large AI model scMulan with 368 million parameters to generate cell transcriptomics with designated attributes by learning the cell language. We defined a unified c-sentence to incorporate cell transcriptomics and meta-attributes, and pre-trained scMulan on the equivalence of 100 million human cells. Experiments showed that scMulan can generate designated pseudo transcriptomics, predict missing attributes of cells, reconstruct unobserved cells along functional gradients, and can help to identify driving regulators of cell fates. The generated data passed tests of current tools and can reflect the underlying biology.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bian, H., Chen, Y., Dong, X., Li, C., Hao, M., Chen, S., Hu, J., Sun, M., Wei, L., Zhang, X.. 2024-01-29. scMulan: a multitask generative pre-trained language model for single-cell analysis. https://doi.org/10.1101/2024.01.25.577152
Cite the original work for its findings. Save a collection to share your selection of sources.