Search bioRxiv⌕ Search

Biology subjects

Strudwick, J.

Publications and source records attributed to Strudwick, J..

2 recordsLinked to original sources

RNA foundation models enable generalizable endometriosis disease classification and stable gene-level interpretation

Endometriosis is a chronic inflammatory condition with significant diagnostic delays impacting one in ten reproductive age women worldwide. While machine learning (ML) models trained on transcriptomic data show promise for disease prediction, limited generalizability across independent patient cohorts has hindered clinical translation. Foundations models (FMs) pretrained on large-scale transcriptomic data offer promise to learn transferrable, biologically meaningful representations that could support cross-cohort predictions. We assembled a 12-cohort bulk RNA-seq benchmark (334 samples) and developed a computationally efficient pipeline to test whether FMs improve endometriosis classification, an approach not previously applied to this disease. Using AutoXAI4Omics with cohort-aware validation, we compared embeddings derived from five state-of-the-art RNA FMs against TPM baselines. In cross-cohort prediction, FM embeddings significantly improved performance, achieving a weighted F1-score of 0.83 vs. 0.68 for the baseline. To allow gene-level interpretation of FM embedding models, we introduce classified-aligned integrated gradients (CA-IG), an interpretability approach aligning gene-level attributions to the downstream classifier without end-to-end finetuning. CA-IG revealed a conserved set of predictive genes from FM embeddings across cohort-validation regimes, contrasting with unstable baseline explainability, suggesting that FM embeddings prioritized transferable disease-related signal over cohort-specific effects. These genes include novel candidates that converge on biologically plausible pathways for endometriosis.

bioinformatics↗

AutoXAI4Omics: an Automated Explainable AI tool for Omics and tabular data

Machine learning (ML) methods have the potential of detailed insights of complex biological systems and today are increasingly used to analyse omics data for tasks such as the discovery of novel biomarkers and phenotype prediction. It can be extremely beneficial and powerful for scientists, domain experts, to easily run sophisticated, robust, and interpretable ML pipelines without the need for an in depth understanding of the code needed to train, tune, optimise ML algorithms. They can then focus on the biological interpretation and validation of the results and insights generated by ML models. Here, we present an entirely automated open-source explainable AI tool, AutoXAI4Omics, that performs classification and regression tasks from omics and tabular numerical data. AutoXAI4Omics accelerates scientific discovery by automating processes and decisions made by AI experts, e.g., selection of the best feature set, hyper-tuning of different ML algorithms and selection of the best ML model for a specific task and dataset. Prior to ML analysis AutoXAI4Omics incorporates feature filtering options that are tailored to specific omic data types. Moreover, the insights into the predictions that are provided by the tool through explainability analysis highlight associations between omic feature values and the targets under investigation e.g., predicted phenotypes, facilitating the discovery of actionable insights. AutoXAI4Omics is at: https://github.com/IBM/AutoXAI4Omics. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=188 HEIGHT=200 SRC="FIGDIR/small/586460v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@366327org.highwire.dtl.DTLVardef@a7b559org.highwire.dtl.DTLVardef@7319aforg.highwire.dtl.DTLVardef@9b5030_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗