Search bioRxiv⌕ Search

Biology subjects

Petrick, L.

Publications and source records attributed to Petrick, L..

4 recordsLinked to original sources

Stability-driven multi-omics integration for reproducible latent structure

High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We propose a Stability-driven framework for multi-omics integration that combines sparse generalized canonical correlation analysis with repeated cross-validation, out-of-sample projection, and systematic evaluation of both component-level and feature-level stability. We apply this framework to untargeted metabolomic and Olink targeted inflammation proteomic profiles in a thyroid cancer case-control cohort (n = 162). Our Stability-driven integration identified reproducible metabolomic and proteomic latent components that showed consistent out-of-sample disease associations and tracked temporally structured changes relative to time to diagnosis. The proposed framework provides a generalizable strategy for identifying reproducible latent structures that improve robustness of biological inference in multi-omics studies.

bioinformatics↗

High-dimensional mediation analysis to elucidate the role of metabolites in the association between PFAS exposure and reduced SARS-CoV-2 IgG in pregnancy

We previously found that per- and polyfluoroalkyl substances (PFAS) mixture exposure is inversely associated with SARS-CoV-2 IgG (IgG) antibody levels in pregnant individuals. Here, we aim to identify metabolites mediating this relationship to elucidate the underlying biological pathways. This cross-sectional study included 59 pregnant participants from a US-based pregnancy cohort. Untargeted metabolomic profiling was performed using Liquid Chromatography-High Resolution Sass spectrometry (LC-HRMS), and weighted Quantile Sum (WQS) regression was applied to assess the PFAS and metabolites mixture effects on IgG. Metabolite indices positively or negatively associated with IgG levels were constructed separately and their mediation effects were examined independently and jointly. The PFAS- index was negatively associated with IgG levels (beta=-0.315, p<0.001), with PFHpS and PFHxS as major contributors. Two metabolite-indices were constructed, one positively (beta=1.249, p<0.001) and one negatively (beta=-1.200, p<0.001) associated with IgG. Key contributors for these indices included trigonelline, adipate, p-octopamine, and n-acetylproline. Analysis of a single mediator showed that 74.6% (95% CI: 45.9%, 98.0%) and 68.6% (95% CI: 41.8%, 94.1%) of the PFAS index-IgG total effect were mediated by the negative and positive metabolites-indices, respectively. Joint analysis of the metabolite-indices indicated a cumulative mediation effect of 83.8% (95% CI: 58.1%, 98.7%). Enriched pathways associated with these metabolites indices were phenylalanine, tyrosine, and tryptophan biosynthesis and arginine metabolism. We observed significant mediation effects of plasma metabolites on the PFAS-IgG relationship, suggesting that PFAS is associated with alteration in the balance of plasma metabolites that contributes to reduced plasma IgG production.

bioinformatics↗

Reactomics: Using mass spectrometry as a chemical reaction detector

Untargeted metabolomics analysis captures chemical reactions among small molecules. Common mass spectrometry-based metabolomics workflows first identify the small molecules significantly associated with the outcome of interest, then begin exploring their biochemical relationships to understand biological fate (environmental studies) or biological impact (physiological response). We suggest an alternative by which biochemical relationships can be directly retrieved through untargeted high-resolution paired mass distance (PMD) analysis without a priori knowledge of the identities of participating compounds. Retrieval is done using high resolution mass spectrometry as a chemical reaction detector, where PMDs calculated from the mass spectrometry data are linked to biochemical reactions obtained via data mining of small molecule and reaction databases, i.e. Reactomics. We demonstrate applications of reactomics including PMD network analysis, source appointment of unknown compounds, and biomarker reaction discovery as a complement to compound discovery analyses used in traditional untargeted workflows. An R implementation of reactomics analysis and the reaction/PMD databases is available as the pmd package (https://yufree.github.io/pmd/).

bioinformatics↗

Data-adaptive pipeline for filtering and normalizing metabolomics data.

IntroductionUntargeted metabolomics datasets contain large proportions of uninformative features and are affected by a variety of nuisance technical effects that can bias subsequent statistical analyses. Thus, there is a need for versatile and data-adaptive methods for filtering and normalizing data prior to investigating the underlying biological phenomena.\n\nObjectivesHere, we propose and evaluate a data-adaptive pipeline for metabolomics data that are generated by liquid chromatography-mass spectrometry platforms.\n\nMethodsOur data-adaptive pipeline includes novel methods for filtering features based on blank samples, proportions of missing values, and estimated intra-class correlation coefficients. It also incorporates a variant of k-nearest-neighbor imputation of missing values. Finally, we adapted an RNA-Seq approach and R package, scone, to select an appropriate normalization scheme for removing unwanted variation from metabolomics datasets.\n\nResultsUsing two metabolomics datasets that were generated in our laboratory from samples of human blood serum and neonatal blood spots, we compared our data-adaptive pipeline with a traditional filtering and normalization scheme. The data-adaptive approach outperformed the traditional pipeline in almost all metrics related to removal of unwanted variation and maintenance of biologically relevant signatures. The R code for running the data-adaptive pipeline is provided with an example dataset at https://github.com/courtneyschiffman/Data-adaptive-metabolomics.\n\nConclusionOur proposed data-adaptive pipeline is intuitive and effectively reduces technical noise from untargeted metabolomics datasets. It is particularly relevant for interrogation of biological phenomena in data derived from complex matrices associated with biospecimens.

bioinformatics↗