Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.08.10.743942

Cancer Cell Line Heterogeneity Imposes a Primary Bottleneck for Virtual Perturbation Screening at Scale

Abstract

UniPert-G2CP (Li et al., Cell, 2026) bridges genetic and chemical screens from molecular representation to phenotype modeling across five cancer cell lines [1]. Here we extend this architecture to 162 cell lines, 32,039 compounds, and genome-wide output (12,328 genes), and find that the resulting prediction platform reveals a striking performance divide: normal and primary cell lines achieve substantially higher prediction fidelity (mean per-cell-line Pearson Correlation Coefficient (PCC) = 0.480, median 0.502) than cancer lines (0.296, median 0.324; Mann-Whitney p = 3.93e-11), identifying cancer cell line heterogeneity as a primary bottleneck for virtual perturbation screening at scale. To understand the sources of this divide, we analyzed per-cell-line performance, genome-wide directional accuracy, mechanism-clustering (SMD), and compound-protein interaction (CPI) enrichment. Directional accuracy on top-5% effect-size genes reaches 73.8% (genome-wide 60.8%), with pathway-dependent recovery: of four literature-supported perturbation-gene pairs queried across three cell lines, two were recapitulated (dexamethasone-TSC22D3/NFKBIA/FKBP5; bortezomib-BAG3/DNAJB1/HSPA1A), one was absent (CD36 depletion-PPARG/CEBPA in ASC), and one was partially recapitulated (metformin-SLC7A5 in HEPG2). Mechanism-clustering SMD of the learned embedding reached 1.636 (vs. original 1.85; 88.5% retention at 32x cell-line coverage), exceeding the ECFP4 fingerprint baseline (1.613), while self-consistency Mantel rho=0.852 confirmed the model retains compound mechanism structure internally. Overall held-out performance: genetic perturbation PCC=0.442 (978-gene subset); novel drug PCC=0.3047 (genome-wide). Analysis of CPI enrichment reveals that training-data overlap inflates apparent performance: 54.9% of Touchstone evaluation pairs overlap with our ChEMBL-derived training CPI pairs, reducing effective EF from 139 to 109 at top 0.5% yet remaining far above random (1.0). These findings establish the first large-scale characterization of cell-type-dependent generalization in perturbation-to-phenotype prediction. The observed performance stratification between normal and cancer lines generates testable hypotheses for why virtual cell models degrade on heterogeneous cancer contexts, and provides a diagnostic framework for identifying where and why such models fail--informing future architecture improvements targeting the CPI vocabulary gap, protein encoder design, and cell-type-aware training strategies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei, K., Zhan, L., Qi, C.. 2026-08-18. Cancer Cell Line Heterogeneity Imposes a Primary Bottleneck for Virtual Perturbation Screening at Scale. https://doi.org/10.64898/2026.08.10.743942

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Senescence-associated KRAS upregulation in peripheral T cells links to premature coronary artery disease

Aims: Premature coronary artery disease (PCAD) lacks specific molecular drivers, and the role of immunosenescence is unclear. We investigated whether aging-related gene dysregulation in T cells contributes to PCAD. Methods: We combined bulk transcriptomics of PBMCs from 12 PCAD patients and 21 controls, single-cell RNA sequencing of PBMCs and human atherosclerotic plaques, weighted gene co-expression network analysis, gene perturbation network analysis, and molecular docking. Results: KRAS was identified as a hub gene intersecting PCAD-associated genes and aging-related genes. Single-cell analysis showed KRAS upregulation predominantly in effector CD8+ T cells, which exhibited the highest senescence scores that were further elevated in disease. Network perturbation of KRAS strongly impacted the cell killing pathway. KRAS-high effector CD8+ T cells were detected in coronary and carotid plaques, displaying enhanced cytotoxicity, exhaustion, and senescence features. Additionally, a candidate small molecule was computationally predicted to bind inactive KRAS. Conclusions: Elevated KRAS expression in senescent, cytotoxic CD8+ T cells is associated with PCAD, bridging immunosenescence and premature atherosclerosis. This finding provides a novel biomarker candidate and potential therapeutic entry point, awaiting further functional validation.

bioinformatics↗

Targeted finetuning enables co-folding models to learn ligand-induced protein conformational states

Advances in protein structure prediction have enabled all-atom protein-ligand co-folding models that predict bound conformations directly from sequence and small-molecule structure. However, these models often fail to generalize to novel binding sites or alternative protein conformational states, limiting their utility for chemical biology and drug discovery. Here we show this limitation reflects training data bias rather than architectural constraints and can be overcome through targeted finetuning. Using ten previously unseen X-ray structures of Werner (WRN) helicase from a drug discovery program, we finetune Boltz-1 to learn both an allosteric binding site and a large conformational change locking the enzyme in an inactive state, while preserving accuracy on the ATP-bound state. The finetuned model generalizes to different chemical series and transfers the conformational logic across RecQ-family helicases in a binding-site sequence-dependent manner. This approach provides a blueprint for adapting foundation models as new structural and mechanistic data emerge, enabling co-folding networks to capture ligand-induced conformational switches and binding poses absent from their training data but central to biological regulation and therapeutic intervention.

bioinformatics↗

Benchmarking single-cell foundation models for aging biology

Single cell foundation models (scFMs) provide representations of cellular states, but their utility across biological questions in aging research remains unclear. We established a benchmark of cellular representations for aging research, evaluating ten general-purpose scFMs, three aging-specific models and conventional methods across five biological questions using more than 2.5 million single cell transcriptomes. Using frozen pretrained representations, Geneformer performed best among scFMs for chronological age prediction and age pseudotime concordance, although 2,000 highly variable genes achieved higher mean performance. Several scFMs captured positive molecular age shifts across three disease contexts, consistent with reported aging-associated changes. SCimilarity performed well for rare cellular state identification across out-of-distribution datasets, exceeding aging specific models and conventional baselines. At the gene level, scGPT showed the highest recovery of reference TF target interactions, including aging-related regulatory hubs. Overall, scFMs supported diverse aging analyses, but performance depended on the biological question, highlighting their utility for rare cellular state identification and regulatory analysis.

bioinformatics↗