bioRxiv · 10.64898/2026.06.23.734145
Learning Perturbation Effects Through Contrastive Alignment of Multimodal Biological Embeddings
Abstract
Multimodal single-cell perturbation screens offer a scalable approach for characterizing the effects of genetic and chemical interventions on cellular state. However, most existing representation-learning methods are tailored to a single perturbation modality and fail to explicitly incorporate external semantic knowledge, which limits their ability to generalize across datasets and perturbation types. Here, we introduce PertOmni, a CLIP-style multimodal representation-learning framework that aligns transcriptomic perturbation signatures with text-derived embeddings of curated genes and compound descriptions, as well as image-derived embeddings from cell paintings. PertOmni jointly trains a shared transcriptomic encoder and dataset-specific text encoders using a masked contrastive objective that emphasizes within-cell-type discrimination while mitigating confounding effects arising from cell-type heterogeneity. We evaluate the produced joint embedding space on bi-directional retrieval, drug-gene interaction inference, and perturbation prediction across both small-molecule and CRISPRi perturbation datasets, and demonstrate consistent improvements over strong baseline methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Long, W., Liu, T., Szalata, A., Theis, F. J., Xue, L., Zhao, H.. 2026-06-26. Learning Perturbation Effects Through Contrastive Alignment of Multimodal Biological Embeddings. https://doi.org/10.64898/2026.06.23.734145
Cite the original work for its findings. Save a collection to share your selection of sources.