Search bioRxiv⌕ Search

Biology subjects

Ancajas, C. M. F.

Publications and source records attributed to Ancajas, C. M. F..

3 recordsLinked to original sources

AI-assisted isolation of bioactive Dipyrimicins from Amycolatopsis azurea and identification of its corresponding dip biosynthetic gene cluster

One of the major challenges in natural product discovery is the prioritization of compounds with useful activities from microbial sources. In particular, this is a challenge in genome mining for novel natural products, where the structures and activities of compounds produced by bioinformatically identified and uncharacterized biosynthetic gene clusters remain unknown. Here, we utilize a machine learning model to predict the antibacterial activity of a natural product from its biosynthetic gene cluster (BGC). We prioritized the strain Amycolatopsis azurea DSM 43854 which was predicted by machine learning to have the capacity to produce multiple natural products with antibacterial activity. Together with bioactivity-guided fractionation, we isolated dipyrimicins A and B from Amycolatopsis azurea DSM 43854 and, for the first time, linked them to their BGC. This dip BGC was predicted by our model to encode a product with 75% antibacterial probability and shares only 40-52% similarity with previously characterized BGCs. We confirmed the antimicrobial properties of the dipyrimicins against a few test strains and identified key tailoring enzymes, including an O-methyltransferase and amidotransferase, that differentiated them from other related 2,2-bipyridine biosynthetic pathways. Importantly, As the dip BGC was not in the training set of the model, our results demonstrate the ability of the model to generalize beyond its training set and the potential of machine learning to accelerate novel bioactive natural product discovery and deorphanization of biosynthetic gene clusters.

biochemistry↗

Assessing the ability of ChatGPT to extract natural product bioactivity and biosynthesis data from publications

Natural products are an excellent source of therapeutics and are often discovered through the process of genome mining, where genomes are analyzed by bioinformatic tools to determine if they have the biosynthetic capacity to produce novel or active compounds. Recently, several tools have been reported for predicting natural product bioactivities from the sequence of the biosynthetic gene clusters that produce them. These tools have the potential to accelerate the rate of natural product drug discovery by enabling the prioritization of novel biosynthetic gene clusters that are more likely to produce compounds with therapeutically relevant bioactivities. However, these tools are severely limited by a lack of training data, specifically data pairing biosynthetic gene clusters with activity labels for their products. There are many reports of natural product biosynthetic gene clusters and bioactivities in the literature that are not included in existing databases. Manual curation of these data is time consuming and inefficient. Recent developments in large language models and the chatbot interfaces built on top of them have enabled automatic data extraction from text, including scientific publications. We investigated how accurate ChatGPT is at extracting the necessary data for training models that predict natural product activity from biosynthetic gene clusters. We found that ChatGPT did well at determining if a paper described discovery of a natural product and extracting information about the products bioactivity. ChatGPT did not perform as well at extracting accession numbers for the biosynthetic gene cluster or producers genome although using an altered prompt improved accuracy.

bioinformatics↗

Statistical Coupling Analysis Predicts Correlated Motions in Dihydrofolate Reductase

The role of dynamics in enzymatic function is a highly debated topic. Dihydrofolate reductase (DHFR), due to its universality and the depth with which it has been studied, is a model system in this debate. Myriad previous works have identified networks of residues in positions near to and remote from the active site that are involved in dynamics and others that are important for catalysis. For example, specific mutations on the Met20 loop in E. coli DHFR (N23PP/S148A) are known to disrupt millisecond-timescale motions and reduce catalytic activity. However, how and if networks of dynamically coupled residues influence the evolution of DHFR is still an unanswered question. In this study, we first identify, by statistical coupling analysis and molecular dynamic simulations, a network of coevolving residues, which possess increased correlated motions. We then go on to show that allosteric communication in this network is selectively knocked down in N23PP/S148A mutant E. coli DHFR. Finally, we identify two sites in the human DHFR sector which may accommodate the Met20 loop double proline mutation while preserving dynamics. These findings strongly implicate protein dynamics as a driving force for evolution.

biochemistry↗