TM-Vec 2: Accelerated Protein Homology Detection for Structural Similarity
Understanding protein function is an essential aspect of many biological applications. The exponential growth of protein sequence data has created a critical throughput bottleneck for structural homology detection: While billions of protein sequences have been identified from DNA sequencing data, the number of protein folds underlying biology is surprisingly limited, likely numbering tens of thousands. The "sequence-fold gap" limits the success of functional annotation methods that rely on sequence homology, especially for newly sequenced genomes. TM-Vec is a deep learning architecture that can predict TM-scores as a metric of structural similarity directly from sequence pairs, bypassing true structural alignment. However, the computational demands of its protein language model (PLM) embeddings create a significant bottleneck for large-scale database searches. In this work we present TM-Vec 2s, a highly efficient model created through distillation from a foundational teacher model. Benchmarks on the CATH and SCOPe domains for large-scale database queries showed that TM-Vec 2s is 185x faster than the original TM-Vec, while achieving higher TM-score prediction accuracy. It is even more efficient than the structure-informed search implemented in Foldseek, while showing strong competition in identifying remote homology between protein molecules. TM-Vec 2s significantly expands scalable, efficient exploration of protein structural homology space.