bioRxiv · 10.1101/2024.12.17.627486
HieVi: Protein Large Language Model for proteome-based phage clustering
Abstract
Viral taxonomy is a challenging task due to the propensity of viruses for recombination. Recent updates from the ICTV and advancements in proteome-based clustering tools highlight the need for a unified framework to organize bacteriophages (phages) across multiscale taxonomic ranks, extending beyond genome-based clustering. Meanwhile, self-supervised large language models, trained on amino acid sequences, have proven effective in capturing the structural, functional, and evolutionary properties of proteins. Building on these advancements, we introduce HieVi, which uses embeddings from a protein language model to define a vector representation of phages and generate a hierarchical tree of phages. Using the INPHARED dataset of 24,362 complete and annotated viral genomes, we show that in HieVi, a multi-scale taxonomic ranking emerges that aligns well with current ICTV taxonomy. We propose that this method, unique in its integration of protein language models for viral taxonomy, can encode phylogenetic relationships, at least up to the family level. It therefore offers a valuable tool for biologists to discover and define new phage families while unraveling novel evolutionary connections.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Panigrahi, S., Ansaldi, M., Ginet, N.. 2024-12-20. HieVi: Protein Large Language Model for proteome-based phage clustering. https://doi.org/10.1101/2024.12.17.627486
Cite the original work for its findings. Save a collection to share your selection of sources.