Search bioRxiv⌕ Search

Biology subjects

Zachariah, D.

Publications and source records attributed to Zachariah, D..

2 recordsLinked to original sources

Predicting allergy and postpartum depression from incomplete compositional microbiome

Time series of compositional data are a common format for many high-throughput studies of biological molecules, e.g., analyzing the response to a treatment or with the aim of predicting an outcome. However, data from some time points may be missing, which reduces the size of the complete dataset. We propose a method for binary classification that includes imputation for missing values and logarithmic transformation of compositional data. Imputation approaches entail models that incorporate artificial data alongside true measurements, thereby supplementing the dataset. We consider two datasets from prospective analyses with as-sociated target labels, aiming to improve prediction accuracy. We predict infants food allergies from their gut microbiome with a balanced accuracy of 0.72. We forecast postpartum depression based on gut microbiome data collected during pregnancy, with a balanced accuracy of 0.62. Features extracted from the microbiome time series, specifically ratios of bacterial abundance, are statistically significant indicators of depression.

bioinformatics↗

Error reduction in leukemia machine learning classification with conformal prediction

PurposeRecent advances in machine learning (ML) have led to the development of classifiers that predict molecular subtypes of acute lymphoblastic leukemia (ALL) using RNA sequencing (RNA-seq) data. While these models have shown promising results, they often lack robust performance guarantees. The aim of this study was three-fold: to quantify the uncertainty of these classifiers; to provide prediction sets that control the false negative rate (FNR); and to perform implicit reduction by transforming incorrect predictions into uncertain predictions. MethodsConformal prediction is a distribution-agnostic framework for generating statistically calibrated prediction sets whose size reflects model uncertainty. In this study, we applied an extension called conformal risk control to ALLIUM, an RNA-seq ALL subtype classifier. Leveraging RNA-seq data from 1042 patient samples taken at diagnosis, we developed a multi-class conformal predictor, ALLCoP, which generates statistically guaranteed FNR-controlled prediction sets. ResultsALLCoP was able to create prediction sets with specified FNR tolerances ranging from 7.5-30%. In a validation cohort, ALLCoP successfully reduced the FNR of the ALLIUM classifier from 8.95% to 3.5%. For cases whose subtype was not previously known, the use of ALLCoP was able to reduce the occurrence of empty predictions from 37% to 17%. Notably, up to 34% of the multiple-class prediction sets included the PAX5alt subtype, suggesting that increased prediction set size may reflect secondary aberrations and biological complexity, contributing to classifier uncertainty. Finally, ALLCoP was validated on two additional RNA-seq ALL subtype classifiers, ALLSorts and ALLCatchR. ConclusionOur results highlight the potential of conformal prediction in enhancing the use of oncological RNA-seq subtyping classifiers and also in uncovering additional molecular aberrations of potential clinical importance.

bioinformatics↗