Search bioRxiv⌕ Search

Biology subjects

Chasapi, M. N.

Publications and source records attributed to Chasapi, M. N..

4 recordsLinked to original sources

ArthroVerse: mapping protein family diversity across arthropod-associated microbiomes

Metagenomic studies of arthropod-associated microbiomes have generated vast amounts of sequence data, yet the functional and structural organization of these proteins remains largely unexplored. Here, we present ArthroVerse, the first comprehensive database of protein families derived from arthropod-associated metagenomes. Non-redundant protein families were generated after rigorous filtering, deduplication, and clustering. The protein families were further annotated with microbial taxonomy, host associations, protein structural information, and Carbohydrate-active enzymes (CAZyme) predictions. The resulting dataset integrates both metagenomic and reference genome-derived proteins, enabling systematic exploration of functional diversity, evolutionary relationships, and host-microbe interactions in insect microbiomes. ArthroVerse provides a valuable resource for the study of microbial ecology and arthropod physiology, offering unprecedented insight into the protein landscape of insect-associated microbial communities.

bioinformatics↗

WasteFams: A database of protein families from global wastewater microbiomes

Wastewater surveillance has emerged as a critical tool for global epidemiology, yet the functional diversity of wastewater microbiomes remains poorly characterized at the protein level. Here, we present WasteFams, the first comprehensive database dedicated to the systematic exploration of protein families in wastewater metagenomic and metatranscriptomic studies worldwide. Integrating data from 580 metagenomes, 132 metatranscriptomes, and 1,709 reference genomes, WasteFams catalogs 3,887 non-redundant protein families (containing {succeq}100 members) derived from over 105 million predicted proteins. Each protein family is enriched with multi-layered annotations, including AlphaFold3 structural predictions, taxonomic classifications, and biome-specific metadata. To further expand their functional annotation, we integrated deep genomic context analysis to link protein families to Mobile Genetic Elements (MGEs), Biosynthetic Gene Clusters (BGCs), Antibiotic Resistance Genes (ARGs), and CRISPR elements. Accessible through the EnvoFams portal, WasteFams provides a user-friendly interface featuring advanced search capabilities, sequence and structural similarity tools, and interactive visualization modules. As global initiatives increasingly leverage wastewater for public health and environmental insights, WasteFams can serve as a critical resource for discovering novel microbial functions, monitoring resistance mechanisms, and exploring the biotechnological potential of secondary metabolites within wastewater-engineered ecosystems.

bioinformatics↗

Quadrupling the protein family space with global metagenomics

The known universe of protein families represents only a small fraction of natures molecular diversity. From 40.3 billion sequences across 40,446 metagenomes, 9,540 metatranscriptomes, and 539 million proteins from 167,415 reference genomes, we identified 608,258 previously uncharacterized protein families with [≥]100 members and 6.5 million families with [≥]25 members, none matching known Pfam domains or reference proteins. This effort doubles the known repertoire of large families and quadruples that of smaller families. Integration of AlphaFold2-based structural predictions with gene-neighborhood and taxonomic analyses enables the characterization of previously unannotated proteins, revealing candidates for both novel and known biological functions in understudied microbial lineages and biomes. This expanded repertoire provides insights into microbial adaptation and broadens the molecular toolkit available for biotechnology, highlighting the power of global metagenomics to uncover hidden protein diversity.

bioinformatics↗

metagRoot: A comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71,091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, Hidden Markov Models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org

bioinformatics↗