Search bioRxiv⌕ Search

Biology subjects

Begzati, A.

Publications and source records attributed to Begzati, A..

4 recordsLinked to original sources

Isotope tracing-based metabolite identificationfor mass spectrometry metabolomics

Modern mass spectrometry-based metabolomics is a key technology for biomedicine, enabling discovery and quantification of a wide array of biomolecules critical for human physiology. Yet, only a fraction of human metabolites have been structurally determined, and the majority of features in typical metabolomics data remain unknown. To date, metabolite identification relies largely on comparing MS2 fragmentation patterns against known standards, related compounds or predicted spectra. Here, we propose an orthogonal approach to identification of endogenous metabolites, based on mass isotopomer distributions (MIDs) measured in an isotope-labeled reference material. We introduce a computational measure of pairwise distance between metabolite MIDs that allows identifying novel metabolites by their similarity to previously known peaks. Using cell material labeled with 20 individual 13C tracers, this method identified 62% of all unknown peaks, including previously never seen metabolites. Importantly, MID-based identification is highly complementary to MS2-based methods in that MIDs reflect the biochemical origin of metabolites, and therefore also yields insight into their synthesis pathways, while MS2 spectra mainly reflect structural features. Accordingly, our method performed best for small molecules, while MS2-based identification was stronger on lipids and complex natural products. Among the metabolites discovered was trimethylglycyl-lysine, a novel amino acid derivative that is altered in human muscle tissue after intensive lifestyle treatment. MID-based annotation using isotope-labeled reference materials enables identification of novel endogenous metabolites, extending the reach of mass spectrometry-based metabolomics.

biochemistry↗

Joint linear modeling of transcriptomics and proteomics is predictive of cancer metastasis

A central goal of conducting omics measurements is to understand how molecular features inform higher-order cell- and tissue-level phenotypes. In particular, multi-omics offers insights into how information encoded by the genome is coordinated through biological layers, resulting in functional outputs1. Due to myriad post-transcriptional regulatory processes, the coordination between mRNA and protein cannot be simply reduced to gene-wise correlation. Yet, both modalities have been shown to serve as representations of biological state, and multi-omics integration has been used to improve these representations. Multi-omics approaches typically do not focus on how mRNA and protein features coordinate, but rather use the additional information for improved prediction or feature selection. Here, instead, we showed that standard linear machine learning models provide an understanding of transcriptomic and proteomic coordination in the context of a biological phenotype of interest, in this case cancer metastasis. We find that, in the context of metastasis, a select subset of proteomic features--reflecting a more concentrated signal relative to the broadly distributed transcriptomic signal--offers additional information to that encoded by transcriptomics, as demonstrated by improved model performance when integrating the two modalities and the relative feature importance of proteomics. Top features show a depletion of gene-product overlap across modalities, indicating that the model primarily leverages instances in which the two modalities are providing complementary information with respect to phenotype. However, in instances when both modalities are selected for a given gene product, there is high information consistency that synergistically bolsters phenotype prediction. Altogether, by using model fits that relate both modalities to phenotype, we observe a nuanced coordination of protein and mRNA, in which both modalities tend to provide consistent information about phenotype, yet benefits remain to incorporating a combination of both complementary and reinforcing signals across modalities.

systems biology↗

Pulmonary primary oxysterol and bile acid synthesis as a predictor of outcomesin pulmonary arterial hypertension

Pulmonary arterial hypertension (PAH) is a rare and fatal vascular disease with heterogeneous clinical manifestations. To date, molecular determinants underlying the development of PAH and related outcomes remain poorly understood. Herein, we identify pulmonary primary oxysterol and bile acid synthesis (PPOBAS) as a previously unrecognized pathway central to PAH pathophysiology. Mass spectrometry analysis of 2,756 individuals across five independent studies revealed 51 distinct circulating metabolites that predicted PAH-related mortality and were enriched within the PPOBAS pathway. Across independent single-center PAH studies, PPOBAS pathway metabolites were also associated with multiple cardiopulmonary measures of PAH-specific pathophysiology. Furthermore, PPOBAS metabolites were found to be increased in human and rodent PAH lung tissue and specifically produced by pulmonary endothelial cells, consistent with pulmonary origin. Finally, a poly-metabolite risk score comprising 13 PPOBAS molecules was found to not only predict PAH-related mortality but also outperform current clinical risk scores. This work identifies PPOBAS as specifically altered within PAH and establishes needed prognostic biomarkers for guiding therapy in PAH. One-Sentence SummaryThis work identifies pulmonary primary oxysterol and bile acid synthesis as altered in pulmonary arterial hypertension, thus establishing a new prognostic test for this disease.

systems biology↗

Multiple freeze-thaw cycles lead to a loss of consistency in poly(A)-enriched RNA sequencing

RNA-Seq is ubiquitous, but depending on the study, sub-optimal sample handling may be required, resulting in repeated freeze-thaw cycles. However, little is known about how each cycle impacts downstream analyses, due to a lack of study and known limitations in common RNA quality metrics, e.g., RIN, at quantifying RNA degradation following repeated freeze-thaws. Here we quantify the impact of repeated freeze-thaw on the reliability of downstream RNA-Seq analysis. To do so, we developed a method to estimate the relative noise between technical replicates independently of RIN. Using this approach we inferred the effect of both RIN and the number of freeze-thaw cycles on sample noise. We find that RIN is unable to fully account for the change in sample noise due to freeze-thaw cycles. Additionally, freeze-thaw is detrimental to sample quality and differential expression (DE) reproducibility, approaching zero after three cycles for poly(A)-enriched samples, wherein the inherent 3 bias in read coverage is more exacerbated by freeze-thaw cycles, while ribosome-depleted samples are less affected by freeze-thaws. The use of poly(A)-enrichment for RNA sequencing is pervasive in library preparation of frozen tissue, and thus, it is important during experimental design and data analysis to consider the impact of repeated freeze-thaw cycles on reproducibility. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=103 SRC="FIGDIR/small/020792v2_ufig1.gif" ALT="Figure 1"> View larger version (19K): org.highwire.dtl.DTLVardef@c2c756org.highwire.dtl.DTLVardef@1ad1c03org.highwire.dtl.DTLVardef@a4136org.highwire.dtl.DTLVardef@13f403f_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗