Search bioRxiv⌕ Search

Biology subjects

Balsam, D.

Publications and source records attributed to Balsam, D..

2 recordsLinked to original sources

EVEE: Interpretable variant effect prediction from genomic foundation model embeddings

Scientific foundation models learn high-dimensional representations from diverse data modalities, yet what they encode and how to extract that knowledge remain open questions. Here we show that probing the internal representations of Evo 2, a 7-billion-parameter genomic foundation model, enables accurate and interpretable genetic variant effect prediction. We introduce a covariance-based probe that captures second-order structure from Evo 2 sequence embeddings to predict variant pathogenicity across variant types and functional consequences, matching or exceeding specialized predictors within their domains. To ground these predictions in known biological mechanisms, we train a complementary panel of probes on existing annotations to detect which genomic properties are disrupted by a variant. This categorized evidence is then integrated with each variants genomic context through a language model to generate variant-specific mechanistic hypotheses. Our pathogenicity predictions correlate with experimental measures of variant function, clinical penetrance, and biobank disease associations while the mechanistic hypotheses are consistent with expert reviews, known mechanism classes, and downstream molecular readouts. We release pathogenicity scores, disruption profiles, and contextualized interpretations for 4.2 million variants from the ClinVar database as an open resource through the Evo Variant Effect Explorer (EVEE). More broadly, this structured probing approach offers a general framework for interrogating foundation models across scientific disciplines and grounding their outputs in existing domain concepts.

genomics↗

Genome modeling and design across all domains of life with Evo 2

All of life encodes information with DNA. While tools for sequencing, synthesis, and editing of genomic code have transformed biological research, intelligently composing new biological systems would also require a deep understanding of the immense complexity encoded by genomes. We introduce Evo 2, a biological foundation model trained on 9.3 trillion DNA base pairs from a highly curated genomic atlas spanning all domains of life. We train Evo 2 with 7B and 40B parameters to have an unprecedented 1 million token context window with single-nucleotide resolution. Evo 2 learns from DNA sequence alone to accurately predict the functional impacts of genetic variation--from noncoding pathogenic mutations to clinically significant BRCA1 variants--without task-specific finetuning. Applying mechanistic interpretability analyses, we reveal that Evo 2 autonomously learns a breadth of biological features, including exon-intron boundaries, transcription factor binding sites, protein structural elements, and prophage genomic regions. Beyond its predictive capabilities, Evo 2 generates mitochondrial, prokaryotic, and eukaryotic sequences at genome scale with greater naturalness and coherence than previous methods. Guiding Evo 2 via inference-time search enables controllable generation of epigenomic structure, for which we demonstrate the first inference-time scaling results in biology. We make Evo 2 fully open, including model parameters, training code, inference code, and the OpenGenome2 dataset, to accelerate the exploration and design of biological complexity.

genomics↗