Search bioRxiv⌕ Search

Biology subjects

Rajkumar, P.

Publications and source records attributed to Rajkumar, P..

5 recordsLinked to original sources

A searchable metadata network graph for microbiome metabolomics

Establishing the biological context of microbial metabolites remains a major challenge. We present microbiomeMASST, a metadata-driven network graph that maps metabolites across 467 available datasets with 144,424 mass spectrometry files from humans, animals, and microbial culture systems. MicrobiomeMASST integrates monocultures, synthetic communities, and host-associated samples across multiple body sites and plants. MS/MS spectra can be queried to trace occurrence across hosts, experimental conditions, and interventions, enabling cross-study integration. We demonstrate this framework by contextualizing microbial-conjugated bile acids and interrogating microbiome-mediated drug metabolism. Screening gut bacteria revealed deprolylation of the angiotensin-converting enzyme (ACE) inhibitor prodrug enalapril. Using microbiomeMASST, we traced this metabolite across human cohorts, microbial isolates, environmental samples, and in Gorilla gorilla. Structural modeling and enzymatic assays showed that microbial deprolylation abolishes ACE inhibition, thereby inactivating its therapeutic effect. Together, microbiomeMASST links MS/MS spectra to biological context, converting isolated observations into an interpretable microbiome map for cross-study analysis.

biochemistry↗

Charting the Undiscovered Metabolome with Synthetic Multiplexing

Most molecular features detected in untargeted metabolomics remain uncharacterized due to the limited scope of existing spectral reference libraries. We synthesized >100,000 biologically inspired compounds using multiplexed reactions, of which 91% were absent from existing structural databases, and searched the resulting MS/MS library across >1.7 billion public spectra, increasing annotation rates by 17.4%. This approach revealed previously undescribed exposure-derived metabolites, including ibuprofen-carnitine. Because ibuprofen has been linked to rhabdomyolysis, reduced mitochondrial function, and impaired muscle recovery in carnitine-limited contexts, we investigated the functional relevance of this conjugate. Ibuprofen-carnitine reduced carnitine transport via the OCTN2 transporter, and in a postpartum mouse muscle injury model, ibuprofen delayed muscle repair that could be rescued by carnitine supplementation, with urinary ibuprofen-carnitine:carnitine ratios tracking this effect. These findings support a hypothesis whereby NSAID-carnitine conjugates compete for carnitine transport, impairing energy metabolism and muscle recovery in susceptible individuals. Synthetic multiplexing thus provides a scalable route to annotate the dark metabolome and generate experimentally testable biological hypotheses.

biochemistry↗

The microbiome diversifies N-acyl lipid pools - including short-chain fatty acid-derived compounds

N-acyl lipids are important mediators of several biological processes including immune function and stress response. To enhance the detection of N-acyl lipids with untargeted mass spectrometry-based metabolomics, we created a reference spectral library retrieving N-acyl lipid patterns from 2,700 public datasets, identifying 851 N-acyl lipids that were detected 356,542 times. 777 are not documented in lipid structural databases, with 18% of these derived from short-chain fatty acids and found in the digestive tract and other organs. Their levels varied with diet, microbial colonization, and in people living with diabetes. We used the library to link microbial N-acyl lipids, including histamine and polyamine conjugates, to HIV status and cognitive impairment. This resource will enhance the annotation of these compounds in future studies to further the understanding of their roles in health and disease and highlight the value of large-scale untargeted metabolomics data for metabolite discovery.

bioinformatics↗

To impute or not to impute in untargeted metabolomics - that is the compositional question

Untargeted metabolomics often produce large datasets with missing values, arising from biological or technical factors, which can undermine statistical analyses and lead to biased biological interpretations. Imputation methods, such as k-Nearest Neighbors (kNN) and Random Forest (RF) regression are commonly used but their effects vary depending on the type of missing data e.g. Missing Completely At Random (MCAR) and Missing Not At Random (MNAR). Here, we determined the impacts of degree and type of missing data on the accuracy of kNN and RF imputation using two datasets: a targeted metabolomic dataset with spiked-in standards and an untargeted metabolomic dataset. We also assessed the effect of compositional data approaches (CoDA), such as the centered log-ratio (CLR) transform, on data interpretation, since these methods are increasingly being used in metabolomics. Overall, we found that kNN and RF performed more accurately when the proportion of missing data across samples for a metabolic feature was low. However, these imputations could not handle MNAR data and generated wildly inflated values or imputed values where none should exist. Furthermore, we show that the proportion of missing values had a strong impact on the accuracy of imputation which affected the interpretation of the results. Our results suggest extreme caution should be used with imputation even with modestly levels of missing data or when the type of missingness is unknown.

bioinformatics↗

Empirically establishing drug exposure records directly from untargeted metabolomics data

Despite extensive efforts, extracting information on medication exposure from clinical records remains challenging. To complement this approach, we developed the tandem mass spectrometry (MS/MS) based GNPS Drug Library. This resource integrates MS/MS data for drugs and their metabolites/analogs with controlled vocabularies on exposure sources, pharmacologic classes, therapeutic indications, and mechanisms of action. It enables direct analysis of drug exposure and metabolism from untargeted metabolomics data independent of clinical records. Our library facilitates stratification of individuals in clinical studies based on the empirically detected medications, exemplified by drug-dependent microbiota-derived N-acyl lipid changes in a cohort with human immunodeficiency virus. The GNPS Drug Library holds potential for broader applications in drug discovery and precision medicine.

bioinformatics↗