Search bioRxiv⌕ Search

Biology subjects

Tule, S.

Publications and source records attributed to Tule, S..

3 recordsLinked to original sources

Comprehensive molecular impact mapping of common and rare variants at GWAS loci

Deep learning sequence to function models can predict the molecular effects of genetic variants, but their predictions are limited to the cell types and assays they are trained on. Here we describe DNACipher, a deep learning model that predicts the effects of genetic variants across diverse biological contexts--including those not directly measured. DNACipher takes 196 kb of genome sequences as input and imputes variant effects across 38,582 cell type-assay combinations. DNACipher generates predictions for >7 times as many contexts as Enformer, which allows for better detection of variant effects at expression quantitative trait loci (eQTLs). We also introduce DNACipher Deep Variant Impact Mapping (DVIM), a method to identify variants with molecular effects at genome-wide association study (GWAS) loci. Application of DVIM to type 1 diabetes (T1D) reduced the mean fine-mapping credible set size from 24 to 1.4 variants per signal. DVIM variants had significantly higher fine-mapping posterior probabilities, and their predicted effects were supported by single-nucleus ATAC-seq and luciferase assays. DVIM also detected 6547 rare variants with molecular effects at 96% of GWAS T1D loci, and these were enriched for associations with immune traits. In summary, DNACipher DVIM prioritises common and rare variants at GWAS loci by predicting molecular effects across a broad range of contexts.

genetics↗

Do Protein Language Models Learn Phylogeny?

Deep machine learning demonstrates a capacity to uncover evolutionary relationships directly from protein sequences, in effect internalising notions inherent to classical phylogenetic tree inference. We connect these two paradigms by assessing the capacity of protein-based language models (pLMs) to discern phylogenetic relationships without being explicitly trained to do so. We evaluate ESM2, ProtTrans and MSA-Transformer relative to classical phylogenetic methods, while also considering sequence insertions and deletions (indels) across 114 Pfam datasets. The largest ESM2 model tends to outperform other pLMs (including the multimodal ESM3) by recovering phylogenetic relationships among homologous protein sequences in both low- and high-gap settings. pLMs agree with conventional phylogenetic methods in general, but more so for protein families with fewer implied indels, highlighting indels as a key factor differentiating classical phylogenetics from pLMs. We find that pLMs preferentially capture broader as opposed to finer evolutionary relationships within a specific protein family, where ESM2 has a sweet spot for highly divergent sequences, at remote distance. Less than 10% of neurons are sufficient to broadly recapitulate classical phylogenetic distances; when used in isolation the difference between the paradigms is further diminished. We show these neurons are polysemantic, shared among different homologous families but never fully overlapping. We highlight the potential of ESM2 as a complementary tool for phylogenetic analysis, especially when extending to remote homologs that are difficult to align and imply complex histories of insertions and deletions.

molecular biology↗

Optimal Phylogenetic Reconstruction of Insertion and Deletion Events

Insertions and deletions (indels) influence the genetic code in fundamentally distinct ways from substitutions, significantly impacting gene product structure and function. Despite their influence, the evolutionary history of indels is often neglected in phylogenetic tree inference and ancestral sequence reconstruction, hindering efforts to comprehend biological diversity determinants and engineer variants for medical and industrial applications. We frame determining the optimal history of indel events as a single Mixed-Integer Programming (MIP) problem, across all nodes in a phylogenetic tree adhering to topological constraints, and all sites implied by a given set of aligned, extant sequences. By disentangling the impact on ancestral sequences at each branch point, this approach identifies the minimal indel events that jointly explain the diversity in sequences mapped to the tips of that tree. MIP can recover alternate optimal indel histories, if available. We evaluated MIP for indel inference on a dataset comprising 15 real phylogenetic trees associated with protein families ranging from 165 to 2000 extant sequences, and on 60 synthetic trees at comparable scales of data and reflecting realistic rates of mutation. Across relevant metrics, MIP outperformed alternative parsimony-based approaches and reported the fewest indel events, on par or below their occurrence in synthetic datasets. MIP offers a rational justification for indel patterns in extant sequences; importantly, it uniquely identifies global optima on complex protein data sets without making unrealistic assumptions of independence or evolutionary underpinnings, promising a deeper understanding of molecular evolution and aiding novel protein design.

bioinformatics↗