Search bioRxiv⌕ Search

Biology subjects

Charusanti, P.

Publications and source records attributed to Charusanti, P..

6 recordsLinked to original sources

Toward understanding drivers of specialized metabolism in Actinomycetota: insights from 1432 transcriptomics datasets for 132 strains

Actinomycetota are major sources of specialized metabolites with applications in drug discovery and agriculture, yet much of their biosynthetic potential remains silent under standard laboratory conditions. Limited understanding of the mechanisms controlling biosynthetic gene cluster (BGC) activation constrains metabolite discovery and production. Here, we generated 1432 RNA-seq datasets from 132 Actinomycetota strains grown in eight media to identify patterns associated with BGC expression. On average, strains expressed 44% of their encoded BGCs across the tested conditions, with more expression observed among known BGCs (61%) compared to uncharacterized BGCs (35%). Co-expression analyses revealed frequent associations between BGCs and transporters, transcriptional regulators, and proteins containing DUF397 and DUF742 domains. Targeted overexpression of candidate genes selected from BGC-associated co-expression modules increased metabolite production, with DUF397- and DUF742-containing operons showing the broadest effects by boosting the levels of several different specialized metabolites. Other genes boosted levels in a metabolite-specific manner. Together, our results support a multilayered model of BGC regulation in Actinomycetota in which BGC expression is shaped by medium composition, BGC-specific regulators, and integration of BGCs into broader transcriptional network modules. By connecting BGC expression to specific media and co-expressed genes, this study also provides a resource for selecting growth conditions and engineering specialized metabolism.

microbiology↗

RetroMol: Parsing a shared encoding from natural products and their biosynthetic gene clusters

Natural products such as polyketides and nonribosomal peptides (NRPs) are important sources of bioactive compounds, including many antibiotics. Many of them are assembled by modular enzyme complexes and further modified and diversified by tailoring reactions encoded by biosynthetic gene clusters (BGCs). Although natural products and their coding BGCs describe different data modalities of the same biochemical process, a unified language to jointly describe their biochemistry is lacking. Here we introduce a sequence-based representation of the core biosynthesis of modular natural products, which we call primary sequences, that bridges chemical structures and BGCs. We also present RetroMol, an algorithm that parses either natural product structures or their encoding BGCs into their primary sequences of natural product building blocks. RetroMol allows for similarity scoring between natural products and BGCs, enabling the retrieval of compounds, BGCs, and a combination of the two, based on their biosynthetic similarity. This can, for instance, be used to retrieve biosynthetically similar but structurally dissimilar compounds, or link natural products to candidate coding BGCs in large experimental datasets. We demonstrate the latter by rediscovering the nocardichelin B BGC as a proof of principle. We also exemplify the utility of biosynthetic similarity by showing various pairs of biosynthetically similar compounds with low structural similarity. Together, these results establish primary sequences as a shared biosynthetic encoding for natural product comparison and BGC prioritization.

bioinformatics↗

Using protein language models for pangenome construction

Current pangenome construction methods rely largely on nucleotide or protein sequence alignment, limiting their ability to detect remote orthologs and semantic relations. We introduce a novel method that leverages protein language model embeddings to capture functional and semantic relationships beyond sequence similarity. Our approach employs approximate nearest-neighbor search coupled with a clustering step utilizing HDBSCAN, DBSCAN, or weighted single-linkage clustering with multiple similarity thresholds. The method utilizes GPU acceleration, dynamic batching, and ONNX optimization to scale approximately linearly with the number of proteins, enabling the analysis of datasets containing millions of proteins. We evaluated our approach on a randomly sampled subset of OrthoDB and the CAFA5 dataset, benchmarking it against SCARAP. SCARAP is a recently published tool with similar performance to a variety of other common tools for computing pangenomics. Our benchmarking demonstrates that our method produces more specific clusters than SCARAP across both datasets. SCARAP excelled in term consistency within clusters on the OrthoDB dataset, where labels are inferred with sequence alignment (using MMseqs2). Both methods face a significant degradation in term consistency when transitioning to the experimentally validated CAFA5 dataset, ultimately resulting in similar term consistency scores for both approaches. Crucially, our approach yields superior cluster quality on both datasets and significantly outperforms SCARAP across all metrics of functional consistency and coherence on the experimental CAFA5 dataset. Finally, we demonstrate the methods scalability and utility by characterizing the pangenome of 1,034 Streptomyces genomes. The pipeline is available for use at our GitHub: https://github.com/jakob949/pan_genome

bioinformatics↗

Pangenome mining of the Streptomyces genus redefines their biosynthetic potential

BackgroundStreptomyces is a highly diverse genus known for the production of secondary or specialized metabolites with a wide range of applications in the medical and agricultural industries. Several thousand complete or nearly-complete Streptomyces genome sequences are now available, affording the opportunity to deeply investigate the biosynthetic potential within these organisms and to advance natural product discovery initiatives. ResultWe performed pangenome analysis on 2,371 Streptomyces genomes, including approximately 1,200 complete assemblies. Employing a data-driven approach based on genome similarities, the Streptomyces genus was classified into 7 primary and 42 secondary MASH-clusters, forming the basis for a comprehensive pangenome mining. A refined workflow for grouping biosynthetic gene clusters (BGCs) redefined their diversity across different MASH-clusters. This workflow also reassigned 2,729 known BGC families to only 440 families, a reduction caused by inaccuracies in BGC boundary detections. When the genomic location of BGCs is included in the analysis, a conserved genomic structure (synteny) among BGCs becomes apparent within species and MASH-clusters. This synteny suggests that vertical inheritance is a major factor in the acquisition of new BGCs. ConclusionOur analysis of a genomic dataset at a scale of thousands of genomes refined predictions of BGC diversity using MASH-clusters as a basis for pangenome analysis. The observed conservation in the order of BGCs genomic locations showed that the BGCs are vertically inherited. The presented workflow and the in-depth analysis pave the way for large-scale pangenome investigations and enhance our understanding of the biosynthetic potential of the Streptomyces genus.

systems biology↗

Maramycin, a cytotoxic isoquinolinequinone terpenoid produced through heterologous expression of a bifunctional indole prenyltransferase /tryptophan indole-lyase in S. albidoflavus

Isoquinolinequinones represent an important family of natural alkaloids with profound biological activities. Heterologous expression of a rare bifunctional indole prenyltransferase /tryptophan indole-lyase enzyme from Streptomyces mirabilis P8-A2 in S. albidoflavus J1074 led to the activation of a putative isoquinolinequinone biosynthetic gene cluster and production of a novel isoquinolinequinone alkaloid, named maramycin (1). The structure of maramycin was determined by analysis of spectroscopic (1D/2D NMR) and MS spectrometric data. The prevalence of this bifunctional biosynthetic enzyme was explored and found to be a recent evolutionary event with only a few representatives in Nature. Maramycin exhibited moderate cytotoxicity against human prostate cancer cell lines, LNCaP and C4-2B. The discovery of maramycin (1) enriched the chemical diversity of natural isoquinolinequinones and also provided new insights into crosstalk between the host biosynthetic genes and the heterologous biosynthetic genes in generating new chemical scaffolds.

bioengineering↗

A treasure trove of 1,034 actinomycete genomes

Filamentous Actinobacteria, recently renamed Actinomycetia, are the most prolific source of microbial bioactive natural products. Studies on biosynthetic gene clusters benefit from or require chromosome-level assemblies. Here, we provide DNA sequences from more than 1,000 isolates: 881 complete genomes and 153 near-complete genomes, representing 28 genera and 389 species, including 244 likely novel species. All genomes are from filamentous isolates of the class Actinomycetia from the NBC culture collection. The largest genus is Streptomyces with 886 genomes including 742 complete assemblies. We use this data to show that analysis of complete genomes can bring biological understanding not previously derived from more fragmented sequences or less systematic datasets. We document the central and structured location of core genes and distal location of specialized metabolite biosynthetic gene clusters and duplicate core genes on the linear Streptomyces chromosome, and analyze the content and length of the terminal inverted repeats which are characteristic for Streptomyces. We then analyze the diversity of trans-AT polyketide synthase biosynthetic gene clusters, which encodes the machinery of a biotechnologically highly interesting compound class. These insights have both ecological and biotechnological implications in understanding the importance of high quality genomic resources and the complex role synteny plays in Actinomycetia biology.

microbiology↗