Search bioRxiv⌕ Search

Biology subjects

Alidoust, N.

Publications and source records attributed to Alidoust, N..

3 recordsLinked to original sources

Back to basics: Observed statistics are sufficient to predict drug responses

Predicting how cells, tissues, and patients will respond to a drug, cytokine, or genetic perturbation is central to biological and clinical reasoning. The practical goal is to estimate the analysis-ready readouts that support this reasoning: which cellular responses are context-dependent, which perturbations reveal shared or divergent mechanisms, and which observations should become the basis for the next experiment. Here we introduce Rhaister, a perturbation-response predictor that operates directly on screen-level summary statistics. By measuring just a few perturbations in a new biological context, Rhaister predicts the unmeasured perturbations by learning how response patterns vary across reference contexts. This formulation applies to both fine-grained molecular readouts, such as transcriptional responses from Tahoe-100M or other large perturbation screen, or on phenotypic endpoints. To train and apply Rhaister on pheno-typic endpoints we created Emerald Bay, a purpose-built Tahoe dataset that unifies multi-day cancer drug perturbation, pooled Mosaic tumor contexts, and paired transcriptomic response measurements*. Across these settings, Rhaister matches or exceeds substantially more expensive virtual-cell models, often achieving the highest values possible in evaluation metrics, while training in seconds and running predictions in milliseconds. On Emerald Bay, Rhaister predicts context-specific drug phenotypes from sensitivity measurements alone and improves further when including transcriptomic information. We also introduce Rhaister-O, predicting drug responses in new contexts from baseline expression alone and, to our knowledge, provides the first zero-shot model for this task. Rhaister establishes summary-statistic perturbation modeling as a fast, interpretable framework for predicting biological response across new contexts.

genomics↗

Tahoe-x1: Scaling Perturbation-Trained Single-CellFoundation Models to 3 Billion Parameters

Foundation models have transformed natural language processing and computer vision, yet their potential in single-cell biology--particularly for complex diseases such as cancer-- remains underexplored. We present Tahoe-x1 (Tx1), a family of perturbation-trained single-cell foundation models with up to 3 billion parameters. Tx1 is pretrained on large-scale single-cell transcriptomic datasets, including the Tahoe-100M perturbation compendium, and fine-tuned for cancer-relevant tasks. Through architectural optimizations, data loader refinements, and efficient training strategies, Tx1 achieves 3-30x higher compute efficiency than prior implementations of cell-state models. Tx1 jointly learns representations of genes, cells, and compounds using a masked-expression generative objective that incorporates a drug token, enabling flexible adaptation to diverse downstream applications. We evaluate Tx1 across four key disease-relevant benchmarks: (1) prediction of overall and context-specific gene essentiality, (2) identification of genes contributing to the hallmarks of cancer, (3) cell-type classification, and (4) prediction of perturbation responses in held-out cellular contexts. Tx1 achieves state-of-the-art performance across all tasks. We release pretrained checkpoints, training code, and evaluation workflows to accelerate the development of perturbation-trained single-cell foundation models for applications in precision oncology and beyond. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=52 SRC="FIGDIR/small/683759v1_ufig1.gif" ALT="Figure 1"> View larger version (13K): org.highwire.dtl.DTLVardef@aa33d5org.highwire.dtl.DTLVardef@310a38org.highwire.dtl.DTLVardef@195f204org.highwire.dtl.DTLVardef@143f7e3_HPS_FORMAT_FIGEXP M_FIG C_FIG

systems biology↗

Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular Modeling

Building predictive models of the cell requires systematically mapping how perturbations reshape each cells state, function, and behavior. Here, we present Tahoe-100M, a giga-scale single-cell atlas of 100 million transcriptomic profiles measuring how each of 1,100 small-molecule perturbations impact cells across 50 cancer cell lines. Our high-throughput Mosaic platform, composed of a highly diverse and optimally balanced "cell village", reduces batch effects and enables parallel profiling of thousands of conditions at single-cell resolution at an unprecedented scale. As the largest single-cell dataset to date, Tahoe-100M enables artificial-intelligence (AI)-driven models to learn context-dependent functions, capturing fundamental principles of gene regulation and network dynamics. Although we leverage cancer models and pharmacological compounds to create this resource, Tahoe-100M is fundamentally designed as a broadly applicable perturbation atlas and supports deeper insights into cell biology across multiple tissues and contexts. By publicly releasing this atlas, we aim to accelerate the creation and development of robust AI frameworks for systems biology, ultimately improving our ability to predict and manipulate cellular behaviors across a wide range of applications.

genomics↗