Search bioRxiv⌕ Search

Biology subjects

Schlecht, N.

Publications and source records attributed to Schlecht, N..

3 recordsLinked to original sources

Rule-Based Deconstruction and Reconstruction of Diterpene Libraries: Categorizing Patterns & Unravelling the Structural Landscape

Terpenoids make up the largest class of specialized metabolites with over 180,000 reported compounds currently across all kingdoms of life. Their synthesis accentuates one of natures most choreographed enzymatic and non-reversible chemistries, leading to an extensive range of structural functionality and diversity. Current terpenoid repositories provide a seemingly endless landscape to systematically survey for information regarding structure, sourcing, and synthesis. Efforts here investigate entries for the 20-carbon diterpenoid variants and deconstruct the complex patterns into simple, categorical groups. This deconstruction approach reduces over 60,000 unique diterpenoid structures to less than 1,000 categorical structures. Furthermore, the majority of diterpene entries (over 75%) can be represented by less than 25 core skeletons. Natural diterpenoid abundance was mapped throughout the tree of life and structural diversity was correlated at an atom-and-bond resolution. Additionally, all identified core structures provide guidelines for predicting how diterpene diversity originates via the mechanisms catalyzed by diterpene synthases. Over 95% of diterpenoid structures rely on cyclization. Here a reconstructive approach is reapplied based on known biochemical rules to model the birth of compound diversity. Reconstruction enabled prediction of highly probable synthesis mechanisms for bioactive taxane-relatives, which were discovered over three decades ago. This computational synthesis validates previously identified reaction products and pathways, as well as enables predicting trajectories for synthesizing real and theoretical compounds. This deconstructive and reconstructive approach applied to the diterpene landscape provides modular, flexible, and an easy-to-use toolset for categorically simplifying otherwise complex or hidden patterns. Significance StatementWe take a deconstructive and reconstructive approach to explore the origins of the diterpene landscape. Introduction of a navigational toolset enables users to survey compound libraries in ways formerly uncharted. Their utility demonstrated here, maps out diterpene cyclization routes, critical intermediate waypoints, and guidance for how to arrive at compounds previously off-the-map. Information acquired from these tools may imply the diterpene landscape is vastly unexplored, with the plateau for discovery potentially still out of sight.

bioinformatics↗

A molecular representation system with a common reference frame for natural products pathway discovery and structural diversity tasks.

Researchers have uncovered hundreds of thousands of natural products, many of which contribute to medicine, materials, and agriculture. However, missing knowledge of the biosynthetic pathways to these products hinders their expanded use. Nucleotide sequencing is key in pathway elucidation efforts, and analyses of natural products molecular structures, though seldom discussed explicitly, also play an important role by suggesting hypothetical pathways for testing. Structural analyses are also important in drug discovery, where many molecular representation systems - methods of representing molecular structures in a computer-friendly format - have been developed. Unfortunately, pathway elucidation investigations seldom use these representation systems. This gap is likely because those systems are primarily built to document molecular connectivity and topology, rather than the absolute positions of bonds and atoms in a common reference frame, the latter of which enables chemical structures to be connected with potential underlying biosynthetic steps. Here, we present a unique molecular representation system built around a common reference frame. We tested this system using triterpenoid structures as a case study and explored the systems applications in biosynthesis and structural diversity tasks. The common reference frame system can identify structural regions of high or low variability on the scale of atoms and bonds and enable hierarchical clustering that is closely connected to underlying biosynthesis. Combined with phylogenetic distribution information, the system illuminates distinct sources of structural variability, such as different enzyme families operating in the same pathway. These characteristics outline the potential of common reference frame molecular representation systems to support large-scale pathway elucidation efforts. Significance StatementStudying natural products and their biosynthetic pathways aids in identifying, characterizing, and developing new therapeutics, materials, and biotechnologies. Analyzing chemical structures is key to understanding biosynthesis and such analyses enhance pathway elucidation efforts, but few molecular representation systems have been designed with biosynthesis in mind. This study developed a new molecular representation system using a common reference frame, identifying corresponding atoms and bonds across many chemical structures. This system revealed hotspots and dimensions of variation in chemical structures, distinct overall structural groups, and parallels between molecules structural features and underlying biosynthesis. More widespread use of common reference frame molecular representation systems could hasten pathway elucidation efforts.

biochemistry↗

Chromosome-scale Salvia hispanica L. (Chia) genome assembly reveals rampant Salvia interspecies introgression

Salvia hispanica L. (Chia), a member of the Lamiaceae, is an economically important crop in Mesoamerica, with health benefits associated with its seed fatty acid composition. Chia varieties are distinguished based on seed color including mixed white and black (Chia pinta) and black (Chia negra). To facilitate research on Chia and expand on comparative analyses within the Lamiaceae, we generated a chromosome-scale assembly of a Chia pinta accession and performed comparative genome analyses with a previously published Chia negra genome assembly. The Chia pinta and negra genome sequences were highly similar as shown by a limited number of single nucleotide polymorphisms and extensive shared orthologous gene membership. There is an enrichment of terpene synthases in the Chia pinta genome relative to the Chia negra genome. We sequenced and analyzed the genomes of 20 Chia accessions with differing seed color and geographic origin revealing population structure within S. hispanica and interspecific introgressions of Salvia species. As the genus Salvia is polyphyletic, its evolutionary history remains unclear. Using large-scale synteny analysis within the Lamiaceae and orthologous group membership, we resolved the phylogeny of Salvia species. This study and its collective resources further our understanding of genomic diversity in this food crop and the extent of inter-species hybridizations in Salvia. PLAIN LANGUAGE SUMMARYChia pinta is an economically important crop due to the high fatty acid present in the seeds. There are multiple types of Chia based on the seeds color including mixed which and black (Chia pinta), black (Chia negra), and white (Chia blanca). We generated a genome assembly of Chia pinta and compared it to existing genome assemblies. While the assemblies are highly similar there are key differences in terpene synthase composition between Chia pinta and Chia negra. We also sequenced 20 other Chia accessions with different seed color and geographic origin to determine a population structure within Chia. We generated genomic resources to further our understanding of this food crop. ABBREVIATIONSBGC Biosynthetic gene cluster BUSCO Benchmarking Universal Single Copy Orthologs GO Gene ontology SNP Single nucleotide polymorphism TIR Terminal inverted repeat TPS Terpene synthase WGS Whole genome shotgun

genomics↗