Search bioRxiv⌕ Search

Biology subjects

Muller, C. L.

Publications and source records attributed to Muller, C. L..

3 recordsLinked to original sources

Variational inference for microbiome survey data with application to global ocean data

Linking sequence-derived microbial taxa abundances to host (patho-)physiology or habitat characteristics in a reproducible and interpretable manner has remained a formidable challenge for the analysis of microbiome survey data. Here, we introduce a flexible probabilistic modeling framework, VI-MIDAS (Variational Inference for MIcrobiome survey DAta analysiS), that enables joint estimation of context-dependent drivers and broad patterns of associations of microbial taxon abundances from microbiome survey data. VI-MIDAS comprises mechanisms for direct coupling of taxon abundances with covariates and taxa-specific latent coupling which can incorporate spatio-temporal information and taxon-taxon interactions. We leverage mean-field variational inference for posterior VI-MIDAS model parameter estimation and illustrate model building and analysis using Tara Ocean Expedition survey data. Using VI-MIDAS latent embedding model and tools from network analysis, we show that marine microbial communities can be broadly categorized into five modules, including SAR11-, Nitrosopumilus-, and Alteromondales-dominated communities, each associated with specific environmental and spatiotemporal signatures. VI-MIDAS also finds evidence for largely positive taxon-taxon associations in SAR11 or Rhodospirillales clades, and negative associations with Alteromonadales and Flavobacteriales classes. Our results indicate that VI-MIDAS provides a powerful integrative statistical analysis framework for discovering broad patterns of associations between microbial taxa and context-specific covariate data from microbiome survey data.

bioinformatics↗

MetaIBS: large-scale amplicon-based meta analysis of irritable bowel syndrome

BackgroundIrritable Bowel Syndrome (IBS) is a chronic functional bowel disorder causing abdominal discomfort, as well as transit deregulation with constipation and/or diarrhea. The pathophysiology of IBS is poorly understood and believed to be multifactorial. The role of gut microbiota in IBS has been investigated in several case-control studies, in particular via 16S rRNA amplicon sequencing surveys. These studies, however, have not yet led to a consistent picture of significant changes in gut microbial compositions across health and disease. One key bottleneck is the modest cohort sizes of most individual studies and a high diversity of experimental, bioinformatics, and statistical analysis approaches across studies. ResultsWe address these shortcomings by presenting MetaIBS, an open-access data repository and associated meta-analysis workflow of thirteen 16S rRNA amplicon datasets comprising both fecal matter and sigmoid biopsy samples spanning {bsim}2,500 IBS and healthy individuals. MetaIBS includes a tailored computational framework that (i) enables coherent de novo processing and taxonomic assignments of the raw 16S rRNA amplicon reads across experimental protocols and sequencing technologies, and (ii) statistical workflows for visualization and analysis at different taxonomic ranks and data granularity. Our statistical meta-analysis shows that popular high-level microbiome summary statistics, including Firmicutes/Bacteroidota ratios or diversity indices, are insufficient for reliable discrimination between IBS patients and healthy controls. Fine-grained multi-method differential abundance and classification analysis, however, can identify sets of differentially abundant taxa that replicate across multiple datasets, including Coprococcus eutactus and Alistipes finegoldii. ConclusionsMetaIBS provides a curated and reproducible data and (meta-)analysis resource for amplicon-based IBS research at unprecedented scale. MetaIBS allows assessing the heterogeneity of IBS cohorts across multiple experimental protocols, sample types, and IBS phenotypes. Our framework will likely contribute to more coherent insights into the role of the microbiome in IBS and the discovery of reliable microbial IBS biomarkers for follow-up functional and translational studies.

microbiology↗

Negative Binomial factor regression with application to microbiome data analysis

The human microbiome provides essential physiological functions and helps maintain host homeostasis via the formation of intricate ecological host-microbiome relationships. While it is well established that the lifestyle of the host, dietary preferences, demographic background, and health status can influence microbial community composition and dynamics, robust generalizable associations between specific host-associated factors and specific microbial taxa have remained largely elusive. Here, we propose factor regression models that allow the estimation of structured parsimonious associations between host-related features and amplicon-derived microbial taxa. To account for the overdispersed nature of the amplicon sequencing count data, we propose Negative Binomial reduced rank regression (NB-RRR) and Negative Binomial co-sparse factor regression (NB-FAR). While NB-RRR encodes the underlying dependency among the microbial abundances as outcomes and the host-associated features as predictors through a rank-constrained coefficient matrix, NB-FAR uses a sparse singular value decomposition of the coefficient matrix. The latter approach avoids the notoriously difficult joint parameter estimation by extracting sparse unit-rank components of the coefficient matrix sequentially. To solve the non-convex optimization problems associated with these factor regression models, we present a novel iterative block-wise majorization procedure. Extensive simulation studies and an application to the microbial abundance data from the American Gut Project demonstrate the efficacy of the proposed procedure. In the American Gut Project data, we identify key factors that strongly link dietary habits and host life style to specific microbial families.

microbiology↗