Search bioRxiv⌕ Search

Biology subjects

Ubas, A. A.

Publications and source records attributed to Ubas, A. A..

2 recordsLinked to original sources

Back to basics: Observed statistics are sufficient to predict drug responses

Predicting how cells, tissues, and patients will respond to a drug, cytokine, or genetic perturbation is central to biological and clinical reasoning. The practical goal is to estimate the analysis-ready readouts that support this reasoning: which cellular responses are context-dependent, which perturbations reveal shared or divergent mechanisms, and which observations should become the basis for the next experiment. Here we introduce Rhaister, a perturbation-response predictor that operates directly on screen-level summary statistics. By measuring just a few perturbations in a new biological context, Rhaister predicts the unmeasured perturbations by learning how response patterns vary across reference contexts. This formulation applies to both fine-grained molecular readouts, such as transcriptional responses from Tahoe-100M or other large perturbation screen, or on phenotypic endpoints. To train and apply Rhaister on pheno-typic endpoints we created Emerald Bay, a purpose-built Tahoe dataset that unifies multi-day cancer drug perturbation, pooled Mosaic tumor contexts, and paired transcriptomic response measurements*. Across these settings, Rhaister matches or exceeds substantially more expensive virtual-cell models, often achieving the highest values possible in evaluation metrics, while training in seconds and running predictions in milliseconds. On Emerald Bay, Rhaister predicts context-specific drug phenotypes from sensitivity measurements alone and improves further when including transcriptomic information. We also introduce Rhaister-O, predicting drug responses in new contexts from baseline expression alone and, to our knowledge, provides the first zero-shot model for this task. Rhaister establishes summary-statistic perturbation modeling as a fast, interpretable framework for predicting biological response across new contexts.

genomics↗

Tahoe-100M: A Giga-Scale Single-Cell Perturbation Atlas for Context-Dependent Gene Function and Cellular Modeling

Building predictive models of the cell requires systematically mapping how perturbations reshape each cells state, function, and behavior. Here, we present Tahoe-100M, a giga-scale single-cell atlas of 100 million transcriptomic profiles measuring how each of 1,100 small-molecule perturbations impact cells across 50 cancer cell lines. Our high-throughput Mosaic platform, composed of a highly diverse and optimally balanced "cell village", reduces batch effects and enables parallel profiling of thousands of conditions at single-cell resolution at an unprecedented scale. As the largest single-cell dataset to date, Tahoe-100M enables artificial-intelligence (AI)-driven models to learn context-dependent functions, capturing fundamental principles of gene regulation and network dynamics. Although we leverage cancer models and pharmacological compounds to create this resource, Tahoe-100M is fundamentally designed as a broadly applicable perturbation atlas and supports deeper insights into cell biology across multiple tissues and contexts. By publicly releasing this atlas, we aim to accelerate the creation and development of robust AI frameworks for systems biology, ultimately improving our ability to predict and manipulate cellular behaviors across a wide range of applications.

genomics↗