Search bioRxiv⌕ Search

Biology subjects

Poli, M.

Publications and source records attributed to Poli, M..

3 recordsLinked to original sources

Genome modeling and design across all domains of life with Evo 2

All of life encodes information with DNA. While tools for sequencing, synthesis, and editing of genomic code have transformed biological research, intelligently composing new biological systems would also require a deep understanding of the immense complexity encoded by genomes. We introduce Evo 2, a biological foundation model trained on 9.3 trillion DNA base pairs from a highly curated genomic atlas spanning all domains of life. We train Evo 2 with 7B and 40B parameters to have an unprecedented 1 million token context window with single-nucleotide resolution. Evo 2 learns from DNA sequence alone to accurately predict the functional impacts of genetic variation--from noncoding pathogenic mutations to clinically significant BRCA1 variants--without task-specific finetuning. Applying mechanistic interpretability analyses, we reveal that Evo 2 autonomously learns a breadth of biological features, including exon-intron boundaries, transcription factor binding sites, protein structural elements, and prophage genomic regions. Beyond its predictive capabilities, Evo 2 generates mitochondrial, prokaryotic, and eukaryotic sequences at genome scale with greater naturalness and coherence than previous methods. Guiding Evo 2 via inference-time search enables controllable generation of epigenomic structure, for which we demonstrate the first inference-time scaling results in biology. We make Evo 2 fully open, including model parameters, training code, inference code, and the OpenGenome2 dataset, to accelerate the exploration and design of biological complexity.

genomics↗

Sequence modeling and design from molecular to genome scale with Evo

The genome is a sequence that completely encodes the DNA, RNA, and proteins that orchestrate the function of a whole organism. Advances in machine learning combined with massive datasets of whole genomes could enable a biological foundation model that accelerates the mechanistic understanding and generative design of complex molecular interactions. We report Evo, a genomic foundation model that enables prediction and generation tasks from the molecular to genome scale. Using an architecture based on advances in deep signal processing, we scale Evo to 7 billion parameters with a context length of 131 kilobases (kb) at single-nucleotide, byte resolution. Trained on 2.7M prokaryotic and phage genomes, Evo can generalize across the three fundamental modalities of the central dogma of molecular biology to perform zero-shot function prediction that is competitive with, or outperforms, leading domain-specific language models. Evo also excels at multi-element generation tasks, which we demonstrate by generating synthetic CRISPR-Cas molecular complexes and entire transposable systems for the first time. Using information learned over whole genomes, Evo can also predict gene essentiality at nucleotide resolution and can generate coding-rich sequences up to 650 kb in length, orders of magnitude longer than previous methods. Advances in multi-modal and multiscale learning with Evo provides a promising path toward improving our understanding and control of biology across multiple levels of complexity.

synthetic biology↗

Characterisation and reprogramming of bacteriophage mv4 integrase recombination specificity

Bacteriophage mv4 is a temperate bacterial virus able to integrate its genome at the 3 end of the tRNASER of Lactobacillus delbrueckii subsp. bulgaricus chromosome through site-specific recombination. Previous investigations revealed that the mv4Int/attP/attB recombination module was atypical compared to conventional heterobivalent tyrosine recombinases, such as the paradigmatic Lambdavirus lambda integrase, suggesting alternative recombination mechanism. In vitro recombination assays with random DNA libraries were used to comprehensively delineate the mv4 recombination system. We showed that mv4Int is a 369-aa protein that exhibits all structural hallmarks of integrases from the Tn916 family and interacts cooperatively with its recombination sites. We established that mv4Int distinguishes itself from classical heterobivalent integrases by a greater tolerance to nucleotide variations in attB and core-attP sites. We demonstrated that, upon considering nucleotide degeneracy, the 21-bp core-attP and attB recombination sites share structural similarities with classical heterobivalent integrase systems, with two 7-bp inverted-repeat regions corresponding to mv4Int core-binding sites surrounding a 7-bp strand-exchange region. Furthermore, our study highlighted compositional biases and nucleotide interdependencies within the core-binding regions that exerted a significant influence on the outcomes of recombination events.

molecular biology↗