Search bioRxiv⌕ Search

Biology subjects

Merlotti, A.

Publications and source records attributed to Merlotti, A..

4 recordsLinked to original sources

Time-series sewage metagenomics can separate the seasonal, human-derived and environmental microbial communities, holding promise for source-attributed surveillance

Sewage metagenomics has risen to prominence in urban population surveillance of pathogens and antimicrobial resistance (AMR). Unknown species with similarity to known genomes cause database bias in reference-based metagenomics. To improve surveillance, we designed this study to recover sewage genomes and develop a quantification and correlation workflow for these genomes and AMR over time. We used longitudinal sewage sampling in seven treatment plants from five major European cities to explore the utility of catch-all sequencing of these population-level samples. Using metagenomic assembly methods, we recovered 2,332 metagenome-assembled genomes (MAGs) from prokaryotic species, 1,334 of which were previously undescribed. These genomes account for [~]69% of sequenced DNA and provide insight into sewage microbial dynamics. Rotterdam (Netherlands) and Copenhagen (Denmark) showed strong seasonal microbial community shifts, while Bologna, Rome, (Italy) and Budapest (Hungary) had occasional blooms of Pseudomonas-dominated communities, accounting for up to [~]95% of sample DNA. Seasonal shifts and blooms present challenges for effective sewage surveillance. We find that bacteria of known shared origin, like human gut microbiota, form communities, suggesting the potential for source-attributing novel species and their ARGs through network community analysis. This could significantly improve AMR tracking in urban environments.

microbiology↗

Correlation measures in metagenomic data: the blessing of dimensionality

Microbiome analysis has revolutionized our understanding of various biological processes, spanning human health, epidemiology (including antimicrobial resistance and horizontal gene transfer), as well as environmental and agricultural studies. At the heart of microbiome analysis lies the characterization of microbial communities through the quantification of microbial taxa and their dynamics. In the study of bacterial abundances, it is becoming more relevant to consider their relationship, to embed these data in the framework of network theory, allowing characterization of features like node relevance, pathway and community structure. In this study, we address the primary biases encountered in reconstructing networks through correlation measures, particularly in light of the compositional nature of the data, within-sample diversity, and the presence of a high number of unobserved species. These factors can lead to inaccurate correlation estimates. To tackle these challenges, we employ simulated data to demonstrate how many of these issues can be mitigated by applying typical transformations designed for compositional data. These transformations enable the use of straightforward measures like Pearsons correlation to correctly identify positive and negative relationships among relative abundances, especially in high-dimensional data, without having any need for further corrections. However, some challenges persist, such as addressing data sparsity, as neglecting this aspect can result in an underestimation of negative correlations.

systems biology↗

Covering Hierarchical Dirichlet Mixture Models on binary data to enhance genomic stratifications in Onco-Hematology

Onco-hematological studies are increasingly adopting statistical mixture models to support the advancement of the genetically-driven classification systems for blood cancer. Targeting enhanced patients stratification based on the sole role of molecular biology attracted much interest and contributes to bring personalized medicine closer to reality. In particular, Dirichlet processes have become the preferred method to approach the fit of mixture models. Usually, the multinomial distribution is at the core of such models. However, despite their advanced statistical formalism, these processes are not to be considered black box techniques and a better understanding of their working mechanisms enables to improve their employment and explainability. Focused on genomic data in Acute Myeloid Leukemia, this work unfolds the driving factors and rationale of the Hierarchical Dirichlet Mixture Models of multinomials on binary data. In addition, we introduce a novel approach to perform accurate patients clustering via multinomials based on statistical considerations. The newly reported adoption of the Multivariate Fishers Non-Central Hypergeometric distributions reveals promising results and outperformed the multinomials in clustering both on simulated and real onco-hematological data. Author summaryExplainable models are particularly attractive nowadays since they have the advantage to convince clinicians and patients. In this work we show that a deeper understanding of the Hierarchical Dirichlet Mixture Model, a non-black box method, can lead to better data modelling. In onco-hematology Hierarchical Dirichlet Mixture Models typically help to cluster molecular alterations rather than patients. Here, an intuitive statistical approach is presented to tackle patient classification based on the Hierarchical Dirichlet Mixture Models outcome. Additionally, molecular alterations are usually modelled by Hierarchical Dirichlet Mixture Models as a mixture of multinomial distributions. This work highlights that the alternative Fishers Non-Central Hypergeometric distribution can provide even better results and can give a higher priority to rare molecular alterations for patient classification.

cancer biology↗

Comprehensive analysis of network reconstruction approaches based on correlation in metagenomic data

Microbiome analysis is transforming our understanding of biological processes related to human health, epidemiology (antimicrobial resistance, horizontal gene transfer) environmental and agricultural studies. At the core of microbiome analysis is the description of microbial communities based on quantification of microbial taxa and dynamics. In the study of bacterial abundances, it is becoming more relevant to consider their relationship, to embed these data in the framework of network theory, allowing characterization of features like node relevance, pathway and community structure. In this work we characterize the principal biases in reconstructing networks from correlation measures, associated with the compositional character of relative abundance data, the diversity of abundances and the presence of unobserved species within a single sample, that might lead to wrong correlation estimates. We show how most of these problems can be overcome by applying typical transformations for compositional data, that allow the application of simple measures such as Pearsons correlation to correctly identify the positive and negative relationships between relative abundances, when data dimensionality is sufficiently high. Some issues remain, like the role of data sparsity, that if not properly addressed can lead to imbalances in correlation coefficient distribution.

systems biology↗