Search bioRxiv⌕ Search

Biology subjects

Cediel-Becerra, J. D. D.

Publications and source records attributed to Cediel-Becerra, J. D. D..

3 recordsLinked to original sources

Integrating targeted genome mining and structure-guided modeling reveals unexplored 7-deazapurine-containing pathways

7-deazapurines are nucleoside analogs that play key roles in nucleic acid modification and can serve as building blocks for diverse, bioactive secondary metabolites. Despite their biological significance, their biosynthetic diversity, distribution, and enzymatic determinants of structural diversification remain poorly understood. Here, we leverage large-scale targeted genome mining, phylogenetic, and network analysis to explore 7-deazapurine-containing pathways across [~]2 million bacterial genomes. We identified over 900 candidate biosynthetic gene clusters (BGCs), grouped into more than 100 families, most of which remain uncharacterized. These GATOR-GC-predicted BGCs were predominantly found in Streptomyces. We then examined enzyme-substrate interactions in three representative pathways: (i) peptidyl-deazapurines, (ii) huimycin, and (iii) dapiramicin A. Molecular docking and molecular dynamics (MD) simulations recapitulated known enzyme-substrate interactions and highlighted candidate catalytic residues governing amide bond formation, methylation, and glycosylation. Using this genome- and structure-guided framework, we identified a candidate BGC for dapiramicin A and proposed tailoring steps, including scaffold methylation and deoxy-sugar formation. These findings expand the known diversity of 7-deazapurine-containing BGCs and demonstrate how integrating genome mining with structural modeling can link BGCs to chemical function, providing a foundation for discovering and characterizing 7-deazapurine-containing secondary metabolites. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/718813v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@c00feforg.highwire.dtl.DTLVardef@156468forg.highwire.dtl.DTLVardef@1326e90org.highwire.dtl.DTLVardef@1f8d57b_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Targeted genome mining with GATOR-GC maps the evolutionary landscape of biosynthetic diversity

Gene clusters, groups of physically adjacent genes that work collectively, are pivotal to bacterial fitness and valuable in biotechnology and medicine. While various genome mining tools can identify and characterize gene clusters, they often overlook their evolutionary diversity, a crucial factor in revealing novel cluster functions and applications. To address this gap, we developed GATOR-GC, a targeted genome mining tool that enables comprehensive and flexible exploration of gene clusters in a single execution. We show that GATOR-GC identified a diversity of over 4 million gene clusters similar to experimentally validated biosynthetic gene clusters (BGCs) that other tools fail to detect. To highlight the utility of GATOR-GC, we identified previously uncharacterized co-occurring conserved genes potentially involved in mycosporine-like amino acid biosynthesis and mapped the taxonomic and evolutionary patterns of genomic islands that modify DNA with 7-deazapurines. Additionally, with its proximity-weighted similarity scoring, GATOR-GC successfully differentiated BGCs of the FK-family of metabolites (e.g., rapamycin, FK506/520) according to their chemistries. We anticipate GATOR-GC will be a valuable tool to assess gene cluster diversity for targeted, exploratory, and flexible genome mining. GATOR-GC is available at https://github.com/chevrettelab/gator-gc. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/639861v1_ufig1.gif" ALT="Figure 1"> View larger version (29K): org.highwire.dtl.DTLVardef@6bf054org.highwire.dtl.DTLVardef@6f3eeaorg.highwire.dtl.DTLVardef@18ba201org.highwire.dtl.DTLVardef@3928dd_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

PARAS: high-accuracy machine-learning of substrate specificities in nonribosomal peptide synthetases

Nonribosomal peptides are diverse natural products with important applications in medicine and agriculture. Bacterial and fungal genomes contain thousands of nonribosomal peptide biosynthetic gene clusters (BGCs) of unknown function, providing a promising resource for peptide discovery. Core structural features of such peptides can be inferred by predicting the substrate(s) of adenylation (A) domains in nonribosomal peptide synthetases (NRPSs). However, existing approaches to A domain prediction rely on limited datasets and often struggle with domains selecting large substrates or from less-studied taxa. Here, we systematically curate and computationally analyse 3,653 A domains and present two high-accuracy specificity predictors, PARAS and PARASECT. A type of A domain with unusually high L-tryptophan specificity was identified through the application of PARAS, and intact protein mass spectrometry to the corresponding NRPS showed it to direct the production of tryptopeptin-related metabolites in Streptomyces species. Together, these technologies will accelerate the characterisation of novel NRPSs and their metabolic products. PARAS and PARASECT are available at https://paras.bioinformatics.nl. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=106 SRC="FIGDIR/small/631717v2_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@8e48aaorg.highwire.dtl.DTLVardef@144c78dorg.highwire.dtl.DTLVardef@890c11org.highwire.dtl.DTLVardef@1776de9_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗