Search bioRxivSearch

Biology subjects

Lesniak, N. A.

Publications and source records attributed to Lesniak, N. A..

3 recordsLinked to original sources

Effective application of machine learning to microbiome-based classification problems

Machine learning (ML) modeling of the human microbiome has the potential to identify microbial biomarkers and aid in the diagnosis of many diseases such as inflammatory bowel disease, diabetes, and colorectal cancer. Progress has been made towards developing ML models that predict health outcomes using bacterial abundances, but inconsistent adoption of training and evaluation methods call the validity of these models into question. Furthermore, there appears to be a preference by many researchers to favor increased model complexity over interpretability. To overcome these challenges, we trained seven models that used fecal 16S rRNA sequence data to predict the presence of colonic screen relevant neoplasias (SRNs; n=490 patients, 261 controls and 229 cases). We developed a reusable open-source pipeline to train, validate, and interpret ML models. To show the effect of model selection, we assessed the predictive performance, interpretability, and training time of L2-regularized logistic regression, L1 and L2-regularized support vector machines (SVM) with linear and radial basis function kernels, decision trees, random forest, and gradient boosted trees (XGBoost). The random forest model performed best at detecting SRNs with an AUROC of 0.695 [IQR 0.651-0.739] but was slow to train (83.2 h) and not inherently interpretable. Despite its simplicity, L2-regularized logistic regression followed random forest in predictive performance with an AUROC of 0.680 [IQR 0.625-0.735], trained faster (12 min), and was inherently interpretable. Our analysis highlights the importance of choosing an ML approach based on the goal of the study, as the choice will inform expectations of performance and interpretability. ImportanceDiagnosing diseases using machine learning (ML) is rapidly being adopted in microbiome studies. However, the estimated performance associated with these models is likely over-optimistic. Moreover, there is a trend towards using black box models without a discussion of the difficulty of interpreting such models when trying to identify microbial biomarkers of disease. This work represents a step towards developing more reproducible ML practices in applying ML to microbiome research. We implement a rigorous pipeline and emphasize the importance of selecting ML models that reflect the goal of the study. These concepts are not particular to the study of human health but can also be applied to environmental microbiology studies.

microbiology

The proton pump inhibitor omeprazole does not promote Clostridium difficile colonization in a murine model

Proton pump inhibitor (PPI) use has been associated with microbiota alterations and susceptibility to Clostridioides difficile infections (CDIs) in humans. We assessed how PPI treatment alters the fecal microbiota and whether treatment promotes CDIs in a mouse model. Mice receiving a PPI treatment were gavaged with 40 mg/kg of omeprazole during a 7-day pretreatment phase, the day of C. difficile challenge, and the following 9 days. We found that mice treated with omeprazole were not colonized by C. difficile. When omeprazole treatment was combined with a single clindamycin treatment, one cage of mice remained resistant to C. difficile colonization, while the other cage was colonized. Treating mice with only clindamycin followed by challenge resulted in C. difficile colonization. 16S rRNA gene sequencing analysis revealed that omeprazole had minimal impact on the structure of the murine microbiota throughout the 16 days of omeprazole exposure. These results suggest omeprazole treatment alone is not sufficient to disrupt microbiota resistance to C. difficile infection in mice that are normally resistant in the absence of antibiotic treatment.

microbiology

Fecal short-chain fatty acids are not predictive of colonic tumor status and cannot be predicted based on bacterial community structure

Colonic bacterial populations are thought to have a role in the development of colorectal cancer with some protecting against inflammation and others exacerbating inflammation. Short-chain fatty acids (SCFAs) have been shown to have anti-inflammatory properties and are produced in large quantities by colonic bacteria which produce SCFAs by fermenting fiber. We assessed whether there was an association between fecal SCFA concentrations and the presence of colonic adenomas or carcinomas in a cohort of individuals using 16S rRNA gene and metagenomic shotgun sequence data. We measured the fecal concentrations of acetate, propionate, and butyrate within the cohort and found that there were no significant associations between SCFA concentration and tumor status. When we incorporated these concentrations into random forest classification models trained to differentiate between people with normal colons and those with adenomas or carcinomas, we found that they did not significantly improve the ability of 16S rRNA gene or metagenomic gene sequence-based models to classify individuals. Finally, we generated random forest regression models trained to predict the concentration of each SCFA based on 16S rRNA gene or metagenomic gene sequence data from the same samples. These models performed poorly and were able to explain at most 14% of the observed variation in the SCFA concentrations. These results support the broader epidemiological data that questions the value of fiber consumption for reducing the risks of colorectal cancer. Although other bacterial metabolites may serve as biomarkers to detect adenomas or carcinomas, fecal SCFA concentrations have limited predictive power.\n\nImportanceConsidering colorectal cancer is the third leading cancer-related cause of death within the United States, it is important to detect colorectal tumors early and to prevent the formation of tumors. Short-chain fatty acids (SCFAs) are often used as a surrogate for measuring gut health and for being anti-carcinogenic because of their anti-inflammatory properties. We evaluated the fecal SCFA concentration of a cohort of individuals with varying colonic tumor burden who were previously analyzed to identify microbiome-based biomarkers of tumors. We were unable to find an association between SCFA concentration and tumor burden or use SCFAs to improve our microbiome-based models of classifying people based on their tumor status. Furthermore, we were unable to find an association between the fecal community structure and SCFA concentrations. Our results indicate that the association between fecal SCFAs, the gut microbiome, and tumor burden is weak.

microbiology