bioRxiv · 10.64898/2025.12.21.695754
Quadrupling the protein family space with global metagenomics
Abstract
The known universe of protein families represents only a small fraction of natures molecular diversity. From 40.3 billion sequences across 40,446 metagenomes, 9,540 metatranscriptomes, and 539 million proteins from 167,415 reference genomes, we identified 608,258 previously uncharacterized protein families with [≥]100 members and 6.5 million families with [≥]25 members, none matching known Pfam domains or reference proteins. This effort doubles the known repertoire of large families and quadruples that of smaller families. Integration of AlphaFold2-based structural predictions with gene-neighborhood and taxonomic analyses enables the characterization of previously unannotated proteins, revealing candidates for both novel and known biological functions in understudied microbial lineages and biomes. This expanded repertoire provides insights into microbial adaptation and broadens the molecular toolkit available for biotechnology, highlighting the power of global metagenomics to uncover hidden protein diversity.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aplakidou, E., Baltoumas, F. A., Chasapi, M. N., Lamari, E., Georgakopoulos-Soares, I., Karatzas, E., Iliopoulos, I., Buluc, A., Finn, R. D., Camargo, A. P., Kyrpides, N., Pavlopoulos, G.. 2025-12-23. Quadrupling the protein family space with global metagenomics. https://doi.org/10.64898/2025.12.21.695754
Cite the original work for its findings. Save a collection to share your selection of sources.