Search bioRxiv⌕ Search

Biology subjects

Ellington, C. N.

Publications and source records attributed to Ellington, C. N..

6 recordsLinked to original sources

Uncertainty-Aware Discrete Diffusion Improves Protein Design

Protein inverse folding involves generating amino acid sequences that adopt a specified 3D structure--a key challenge in structural biology and molecular engineering. While discrete diffusion models have demonstrated strong performance, existing methods often apply uniform denoising across residues, overlooking position-specific uncertainty. We propose an uncertainty-aware discrete denoising diffusion model that employs a prior-posterior signaling mechanism to dynamically guide the denoising process. Our approach further integrates learned priors from a pretrained protein large language model and a structure encoder within a modular framework, jointly optimized through multi-objective training. Across multiple benchmarks, our method achieves substantial improvements over state-ofthe-art baselines, offering a principled framework for structure-conditioned sequence generation in proteins and beyond.

bioinformatics↗

Rapid and Reproducible Multimodal Biological Foundation Model Development with AIDO.ModelGenerator

Foundation models (FMs) for DNA, RNA, proteins, cells, and tissues have begun to close long-standing performance gaps in biological prediction tasks, yet each modality is usually studied in isolation. Bridging them requires software that can ingest heterogeneous data, apply large pre-trained backbones from various sources, and perform multimodal benchmarking studies at scale. We present AIDO.ModelGenerator, an open-source toolkit that turns these needs into declarative experiment recipes through a structured experimental framework. AIDO.ModelGenerator provides (i) 300+ datasets covering DNA, RNA, protein, cell, spatial, and multimodal data types; (ii) 30+ pretrained FMs ranging from 3M to 16B parameters; (iii) 10+ plug-and-play use-cases covering inference, adaptation, prediction, generation, and zero-shot evaluation; and (iv) YAML-driven experiment recipes that enable exact reproducibility. On a sequence-to-expression prediction task, AIDO.ModelGenerator systematically builds and tests unimodal and multimodal models, achieving a new SOTA by combining DNA and RNA FMs that outperforms unimodal baselines by over 10%. In a Crohns disease case-study, the frameworks simulated knockout protocol ranks the clinically implicated target SOX4 6,000 positions higher than differential-expression baselines, illustrating its utility for therapeutic target discovery. We release code, tutorials, checkpoints, datasets, and API reference to accelerate multimodal FM research in the life sciences1.

bioinformatics↗

Accurate and General DNA Representations Emerge from Genome Foundation Models at Scale

Language models applied to protein sequences have become a panacea, enabling therapeutics development, materials engineering, and core biology research. Despite the successes of protein language models, genome language models remain nascent. Recent studies suggest the bottleneck is data volume or modeling context size, since long-range interactions are widely acknowledged but sparsely annotated. However, it may be the case that even short DNA sequences are modeled poorly by existing approaches, and current models are unable to represent the wide array of functions encoded by DNA. To study this, we develop AIDO.DNA, a pretrained module for DNA representation in an AI-driven Digital Organism [1]. AIDO.DNA is a seven billion parameter encoder-only transformer trained on 10.6 billion nucleotides from a dataset of 796 species. By scaling model size while maintaining a short context length of 4k nucleotides, AIDO.DNA shows substantial improvements across a breadth of supervised, generative, and zero-shot tasks relevant to functional genomics, synthetic biology, and drug development. Notably, AIDO.DNA outperforms prior encoder-only architectures without new data, suggesting that new scaling laws are needed to achieve computeoptimal DNA language models. Models and code are available through Model-Generator in https://github.com/genbio-ai/AIDO and on Hugging Face at https://huggingface.co/genbio-ai.

bioinformatics↗

Scaling Dense Representations for Single Cell with Transcriptome-Scale Context

Developing a unified model of cellular systems is a canonical challenge in biology. Recently, a wealth of public single-cell RNA sequencing data as well as rapid scaling of self-supervised learning methods have provided new avenues to address this longstanding challenge. However, rapid parameter scaling has been essential to the success of large language models in text and images, while similar scaling has not been attempted with Transformer architectures for cellular modeling. To produce accurate, transferable, and biologically meaningful representations of cellular systems, we develop AIDO.Cell, a pretrained module for representing gene expression and cellular systems in an AI-driven Digital Organism [1]. AIDO.Cell contains a series of 3M, 10M, 100M, and 650M parameter encoder-only dense Transformer models pre-trained on 50 million human cells from diverse tissues using a read-depth-aware masked gene expression pretraining objective. Unlike previous models, AIDO.Cell is capable of handling the entire human transcriptome as input without truncation or sampling tricks, thus learning accurate and general representations of the human cells entire transcriptional context. This pretraining with a longer context was enabled through FlashAttention-2, mixed precision, and large-scale distributed systems training. AIDO.Cell (100M) achieves state-of-the-art results in tasks such as zero-shot clustering, cell-type classification, and perturbation modeling. Our findings reveal interesting loss scaling behaviors as we increase AIDO.Cells parameters from 3M to 650M, providing insights for future directions in single-cell modeling. Models and code are available through ModelGenerator in https://github.com/genbio-ai/AIDO and on Hugging Face.

bioinformatics↗

A Large-Scale Foundation Model for RNA Function and Structure Prediction

Originally marginalized as an intermediate in the information flow from DNA to protein, RNA has become the star of modern biology, holding the key to precision therapeutics, genetic engineering, evolutionary origins, and our understanding of fundamental cellular processes. Yet RNA is as mysterious as it is prolific, serving as an information store, a messenger, and a catalyst, spanning many underchar-acterized functional and structural classes. Deciphering the language of RNA is important not only for a mechanistic understanding of its biological functions but also for accelerating drug design. Toward this goal, we introduce AIDO.RNA, a pre-trained module for RNA in an AI-driven Digital Organism [1]. AIDO.RNA contains a scale of 1.6 billion parameters, trained on 42 million non-coding RNA (ncRNA) sequences at single-nucleotide resolution, and it achieves state-of-the-art performance on a comprehensive set of tasks, including structure prediction, genetic regulation, molecular function across species, and RNA sequence design. AIDO.RNA after domain adaptation learns to model essential parts of protein translation that protein language models, which have received widespread attention in recent years, do not. More broadly, AIDO.RNA hints at the generality of biological sequence modeling and the ability to leverage the central dogma to improve many biomolecular representations. Models and code are available through ModelGenerator in https://github.com/genbio-ai/AIDO and on Hugging Face.

bioinformatics↗

Contextualized Networks Reveal Heterogeneous Transcriptomic Regulation in Tumors at Sample-Specific Resolution

Cancers are shaped by somatic mutations, microenvironment, and patient background, each altering gene expression and regulation in complex ways, resulting in heterogeneous cellular states and dynamics. Inferring gene regulatory networks (GRNs) from expression data can help characterize this regulation-driven heterogeneity, but network inference requires many statistical samples, limiting GRNs to cluster-level analyses that ignore intra-cluster heterogeneity. We propose to move beyond coarse analyses of pre-defined subgroups by using contextualized learning, a multi-task learning paradigm that uses multi-view contexts including phenotypic, molecular, and environmental information to infer personalized models. With sample-specific contexts, contextualization enables sample-specific models and even generalizes at test time to predict network models for entirely unseen contexts. We unify three network model classes (Correlation, Markov, Neighborhood Selection) and estimate context-specific GRNs for 7997 tumors across 25 tumor types, using copy number and driver mutation profiles, tumor microenvironment, and patient demographics as model context. Our generative modeling approach allows us to predict GRNs for unseen tumor types based on a pan-cancer model of how somatic mutations affect gene regulation. Finally, contextualized networks enable GRN-based precision oncology by providing a structured view of expression dynamics at sample-specific resolution, explaining known biomarkers in terms of network-mediated effects and leading to novel subtypings that improve survival prognosis. We provide a SKLearn-style Python package https://contextualized.ml for learning and analyzing contextualized models, as well as interactive plotting tools for pan-cancer data exploration at https://github.com/cnellington/CancerContextualized. Significance StatementNetwork estimation is essential for understanding the structure and function of biological systems, but current statistical approaches fail to capture inter-subject heterogeneity or cross-modality information flow, both of which are needed for understanding complex phenotypes and pathologies. We introduce contextualized network inference, leveraging multi-view contextual metadata to capture similarities and differences among heterogeneous observations during network estimation. Sharing information across contexts enables inference at sample-specific resolution, thus quantifying variation between subjects and revealing context-specific network rewiring. Applied to tumor-specific transcriptional network inference using clinical, molecular, and multi-omic data, contextualized networks improve accuracy, generalize to unseen cancer types, and discover novel prognostic tumor subtypes. By tailoring disease models to each sample, contextualized networks promise to enable precision medicine at unprecedented resolution.

bioinformatics↗