Search bioRxiv⌕ Search

Biology subjects

Yin, W.-B.

Publications and source records attributed to Yin, W.-B..

2 recordsLinked to original sources

SIDERITE: Unveiling Hidden Siderophore Diversity in the Chemical Space Through Digital Exploration

Siderophores, a highly diverse family of secondary metabolites, play a crucial role in facilitating the acquisition of the essential iron. However, the current discovery of siderophore relies largely on manual approaches. In this work, we introduced SIDERTE, a digitized siderophore information database containing 872 siderophore records with 649 unique structures. Leveraging this digitalized dataset, we gained a systematic overview of siderophores by their clustering patterns in the chemical space. Building upon this, we developed a functional group-based method for predicting new iron-binding molecules. Applying this method to 4,314 natural product molecules from TargetMols Natural Product Library for high throughput screening, we experimentally confirmed that 40 out of the 48 molecules predicted as siderophore candidates possessed iron-binding abilities. Expanding our approach to the COCONUT natural product database, we predicted a staggering 3,199 siderophore candidates, showcasing remarkable structure diversity that are largely unexplored. Our study provides a valuable resource for accelerating the discovery of novel iron-binding molecules and advancing our understanding towards siderophores.

biochemistry↗

Non-ribosomal peptide synthetase domain boundary identification and new motifs discovery based on motif-intermotifs standardized architecture

Non-ribosomal peptide synthetase (NRPS) is a diverse family of biosynthetic enzymes for the assembly of bioactive peptides. Despite advances in microbial sequencing, the lack of a consistent standard for annotating NRPS domains and modules has made data-driven discoveries challenging. To address this, we introduced a standardized architecture for NRPS, by using known conserved motifs to partition typical domains. This motif-and-intermotif standardization allowed for systematic evaluations of sequence properties from a large number of NRPS pathways, resulting in the most comprehensive cross-kingdom C domain subtype classifications to date, as well as the discovery and experimental validation of novel conserved motifs with functional significance. Furthermore, our coevolution analysis revealed important barriers associated with reengineering NRPSs and uncovered the entanglement between phylogeny and substrate specificity in NRPS sequences. Our findings provide a comprehensive and statistically insightful analysis of NRPS sequences, opening avenues for future data-driven discoveries. Author SummaryNRPS, a gigantic enzyme that produces diverse microbial secondary metabolites, provides a rich source for important medical products including antibiotics. Despite the extensive knowledge gained about its structure and the large amount of sequencing data available, the frequent failure of reengineering NRPS in synthetic biology highlights the fact that much is still unknown. In this work, we applied existing knowledge to data mining of NRPS sequences, using well-known conserved motifs to partition NRPS sequences into motif-intermotif architectures. This standardization allows for integrating large amounts of sequences from different sources, providing a comprehensive overview of NRPSs across different kingdoms. Our findings included new C domain subtypes, novel conserved motifs with implication in structural flexibility, and insights into why NRPSs are so difficult to reengineer. To facilitate researchers in related fields, we constructed an online platform "NRPS Motif Finder" for parsing the motif-and-intermotif architecture and C domain subtype classification (http://www.bdainformatics.org/page?type=NRPSMotifFinder). We believe that this knowledge-guided approach not only advances our understanding of NRPSs but also provides a useful methodology for data mining in large-scale biological sequences.

bioinformatics↗