Search bioRxiv⌕ Search

Biology subjects

Waman, V.

Publications and source records attributed to Waman, V..

3 recordsLinked to original sources

ContrasTED: contrastive domain embeddings for scalable remote homology classification

Protein structure prediction has expanded structural databases to hundreds of millions of domains. Classifying these domains into homologous superfamilies reveals evolutionary and functional relationships that can persist despite low sequence similarity. As the size of structural databases continues to grow, homology classification requires methods that combine scalability with accuracy. Here we present ContrasTED, which uses CATH-supervised center-contrastive learning to project structure-aware embeddings into a domain-level metric space for nearest-centroid superfamily assignment. On a sequence-filtered S20 benchmark (n = 1,028), superfamily assignment accuracy reached 92.9% (1-NN) and 91.4% (nearest centroid), exceeding sequence search, profile HMMs, Foldseek, and a classifier trained on embeddings. The learned latent space separates superfamilies while retaining structural information below 20% sequence identity, with the largest gains among sparsely represented superfamilies. ContrasTED produces 4.67 million new candidate assignments across 3,796 superfamilies in The Encyclopedia of Domains (TED), extending annotation coverage beyond previous structure-based methods.

bioinformatics↗

Understanding structural and functional diversity of ATP-PPases using protein domains and functional families in CATH database

ATP-Pyrophosphatases (ATP-PPases) are the most primordial lineage of the large and diverse HUP (HIGH-motif proteins, Universal Stress Proteins, ATP-Pyrophosphatase) superfamily. There are four different ATP-PPase substrate-specificity groups, and members of each group show considerable sequence variation across the domains of life despite sharing the same catalytic function. Over the past decade, there has been a >20-fold expansion in the number of ATP-PPase domain structures most recently from advances in protein structure prediction (e.g. Alphafold2). Using the enriched structural information, we have characterised the two most populated ATP-PPase substrate-specificity groups, the NAD-synthases (NAD) and GMP synthases (GMPS). We performed local structural and sequence comparisons between the NADS and GMPS from different domains of life and identified taxonomic-group specific structural functional motifs. As GMPS and NADS are potential drug targets of pathogenic microorganisms including Mycobacterium tuberculosis, structural motifs specific to bacterial GMPS and NADS provide new insights that may aid antibacterial-drug design.

bioinformatics↗

CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models

1.CATH is a protein domain classification resource that combines an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues that might be missed by state-of-the-art HMM-based approaches. The proposed algorithm for this task (CATHe) combines a neural network with sequence representations obtained from protein language models. The employed dataset consisted of remote homologues that had less than 20% sequence identity. The CATHe models trained on 1773 largest, and 50 largest CATH superfamilies had an accuracy of 85.6+-0.4, and 98.15+-0.30 respectively. To examine whether CATHe was able to detect more remote homologues than HMM-based approaches, we employed a dataset consisting of protein regions that had annotations in Pfam, but not in CATH. For this experiment, we used highly reliable CATHe predictions (expected error rate <0.5%), which provided CATH annotations for 4.62 million Pfam domains. For a subset of these domains from homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold structures with experimental structures from the CATHe predicted superfamilies.

bioinformatics↗