Search bioRxiv⌕ Search

Biology subjects

Vollmer, S.

Publications and source records attributed to Vollmer, S..

5 recordsLinked to original sources

Explainable Artificial Intelligence for Cross-Dataset Generalizable Biomarker Discovery in Cardiovascular diseases (CVDs)

CVDs are heterogeneous, multifactorial disorders that remain the leading cause of global mortality from infancy to old age. It requires an early identification and treatment of risk factors to accelerate disease prevention and morbidity improvement. Advancements in transcriptomics technologies gives large pool of heterogenous gene expression data. The technical heterogeneity of gene expression data reduces ability to compare multiple cross-platform datasets at once. To bridge gap, we systematically evaluate three data harmonization techniques: Shambhala-2, TDM, and UPC to align heterogeneous data into a shared expression space while preserving biological signals. Our pipeline integrates 25 independent datasets comprising 983 samples across 23 distinct CVDs phenotypes from both RNA-seq and microarray platforms. The framework benchmarks 35 Machine learning (ML) and Deep learning (DL) classifiers, including Transformers and ResNets, across three data modalities such as RNA-seq, microarray hybridization and RNA-seq + microarray and multiple tissue types. To ensure clinical trustworthiness, we apply multiple Explainable artificial intelligence (XAI) methods, such as SHapley additive exPlanations (SHAP) and Integrated gradientss (IGs), and assess their reliability using quantitative metrics like Area over the perturbation curve (AOPC), Sensitivity, and Infidelity. Results indicate that Shambhala-2 provides superior harmonization by maximizing the biological signal-to-platform ratio. Evaluation of XAI methods reveals that Shapley-based approaches offer the highest stability for identifying influential genomic features in high-dimensional data. Functional enrichment and pathway analyses further confirmed the involvement of identified biomarkers in key cardiovascular processes, including inflammation, immune regulation, oxidative stress, and vascular remodeling. Collectively, this study provides a scalable and interpretable road-map that integrates XAI with cross-dataset biomarker discovery, supporting the transition toward precision cardiology.

bioinformatics↗

Towards Understanding The Relationship Between Brain Disorders and the Gut Microbiome with Explainable Graph Neural Networks

MotivationThe communication between the gut microbiome and the brain, known as the microbiome-gut-brain axis (MGBA), is emerging as a critical factor in neurological and psychiatric disorders. This communication involves complex pathways including neural, hormonal, and immune interactions that enable gut microbes to modulate brain function and behavior. However, the specific mechanisms through which gut microbes influence brain function remain poorly understood, and existing computational efforts to understand these mechanisms are simplistic or have limited scope. ResultsThis work presents a comprehensive approach for elucidating the interactions that allows gut microbes to influence brain disorders. We construct a large curated biomedical knowledge graph comprising 586,318 nodes across 16 entity types and 3,573,936 edges spanning 103 relation types, integrating ontological and experimental data relevant to the MGBA. On this graph, we train GNN-GBA, a GraphSAGE-based graph neural network with a DistMult relation-aware decoder, achieving an AUC-ROC of 0.997 and an F1-score of 0.981 on link prediction, outperforming nine baseline methods across four categories. Using GNNExplainer, we extract and rank mechanistic pathways connecting gut microbes to brain disorders, and demonstrate their stability across multiple random initializations. GNN-GBA successfully identified pathways for 125 brain disorders, revealing shared metabolite hubs (including flavonoids, bile acids, and short-chain fatty acids) that mediate gut-brain communication across diverse neurological conditions. Furthermore, we show that the top pathways are consistent with existing literature for three common disorders. Lastly, we develop an interactive dashboard (GutBrainExplorer) to explore thousands of potential mechanistic pathways across 125 brain disorders, which is publicly available at https://sds-genetic-interaction-analysis.opendfki.de/gut_brain/. AvailabilityCode and data are available at https://github.com/naafey-aamer/GNN-GBA. Contactnaafey.aamer@cs.rptu.de

bioinformatics↗

Artificial Intelligence Powered Biomarker Discovery: A Large-Scale Analysis of 236 Studies Across 19 Therapeutic Areas and 147 Diseases

Biomarkers are the molecular signatures that drive and reflect disease states and are indispensable for disease diagnosis, therapeutic target identification, and drug development. The landscape of biomarker discovery has undergone a transformative shift with the emergence of AI-powered predictive pipelines that can integrate complex, high-dimensional datasets. However, the field still lacks a comprehensive, cross-disciplinary foundation that unites AI pipelines with disease-specific biological insights. Together, a combined scattered knowledge of 15 review articles fails to provide a unified framework encompassing data availability, methodological trends, and disease-specific biomarker discoveries across therapeutic areas. Most prior efforts have concentrated on narrow aspects, either focusing on disease-specific AI models or individual stages of the biomarker discovery pipelines, leaving a substantial gap in translational utility. This study addresses this gap by systematically consolidating and analyzing findings from 236 AI-driven biomarker discovery studies. We systematically map the trends of datasets, data modalities, preprocessing strategies, feature engineering methods, AI models, and explainability methods across 147 diseases, which we further organize into 19 therapeutic areas. By doing so, we aim to provide a comprehensive resource that not only highlights current trends and gaps but also lays the groundwork for future advancements, including the design of multi-task learning models and multimodal AI frameworks tailored to complex biomedical data.

bioinformatics↗

Tick-borne coinfections modulate CD8+ T cell response and progressive leishmaniosis

Leishmania infantum causes human Visceral Leishmaniasis and Leishmaniosis (CanL) in reservoir host, dogs. As infection progresses to disease in both humans and dogs there is a shift from controlling, Type 1, immunity to a regulatory, exhausted, T cell phenotype. In endemic areas, association between tickborne coinfections (TBC) and Leishmania diagnosis and/or clinical severity has been demonstrated. To identify immune factors correlating with disease progression, we prospectively evaluated a cohort of L. infantum infected dogs from 2019-2022. The cohort was TBC-negative with asymptomatic leishmaniosis at the time of enrollment. We measured TBC serology, anti-Leishmania antigen T cell immunity, CanL serological response, parasitemia, and disease severity to probe how nascent TBC perturbs the immune state. At the conclusion, TBC+ dogs with CanL experienced greater increases in anti-Leishmania antibody reactivity and parasite burden compared to dogs that did not have incident TBC during the study. TBC+ dogs were twice as likely to experience moderate (LeishVet stage 2) or severe/terminal disease (LeishVet stage 3/4). Prolonged exposure to TBC was associated with a shift in Leishmania antigen-induced IFN{gamma}/IL-10 and enhanced CD8 T cell proliferation. Frequency of proliferating CD8 T cells significantly correlated with parasitemia and antibody reactivity. TBC exacerbated parasite burden and immune exhaustion. These findings highlight the need for combined vector control efforts as prevention programs for dogs in Leishmania endemic areas to reduce transmission to humans. Public health education efforts should aim to increase awareness of the connection between TBC and leishmaniosis.

immunology↗

FROM CANCER MOLECULAR SUBTYPE TO AI HYPE: BENCHMARKING AI IN CANCER MOLECULAR SUBTYPING

BackgroundCancer molecular subtype classification is an essential component of precision oncology which provides insights into cancer prognosis and guides targeted therapy. Despite the growing applications of AI for cancer molecular subtype classification, challenges persist due to non-standardized dataset configurations, diverse omics modalities, and inconsistent evaluation measures. These issues limit the comparability, reproducibility, and generalizability of AI classifiers across different cancers and hinder the development of robust and accurate AI-driven tools. ResultsThis study benchmarks 35 unique AI classifiers across 153 datasets, covering 8 omics modalities and 20 different cancers. Particularly, it investigates 6 different research questions, and based on comprehensive performance analyses of the 35 AI classifiers it elucidates the research questions with the following answers: (i) Out of 17 different configurations for 5/8 omics modalities, RPPA (RPPA), Gistic2-all-data-by-genes (CNV), HM27 (Meth), and HiSeqV2-exon (Exon) configurations consistently yield better performance; (ii) In terms of 8 omics modalities, RNASeq, miRNA, CNV, and Exon generally achieve higher macro-accuracy compared to Meth., Array, SNP and RPPA; (iii) SNP and RPPA modalities are prone to biases due to technical noise and data imbalance; (iv) Traditional machine learning (ML) models (SVM, XGB, HGB) perform best on small and low-dimensional datasets, while deep learning (DL) models (ResNet18, CNN, NN, MLP) excel on large and high-dimensional datasets; (v) SVM achieves the highest mean macro-accuracy across all classifiers, with NN, ResNet18, DEEPGENE, and MLP also demonstrate strong performance; and (vi) DL classifiers show superior macro accuracy as compared to ML classifiers in 12 out of 20 cancers. ConclusionsThe findings offer key insights to guide the development of standardized, robust, and efficient AI-driven pipelines for cancer molecular subtype classification. This study enhances reproducibility and facilitates better comparison across AI methods, ultimately advancing precision oncology. Key PointsO_LIThis study benchmarks 35 unique AI classifiers, ranging from simpler ML models such as Support Vector Machines (SVM), Histogram-Based Gradient Boosting (HGB), and K-Nearest Neighbors (KNN), to complex DL classifiers including Convolutional Neural Networks (CNNs), computer vision models like DenseNet and ResNet, sequential models such as Recurrent Neural Networks (RNN), Gated Recurrent Units (GRU), Long Short-Term Memory networks (LSTM), and their hybrid combinations (e.g., CNN-LSTM, CNN-GRU), as well as transformer-based models, across 153 datasets spanning 8 omics modalities and 20 cancers. It identifies optimal data configurations and evaluates the performance of these classifiers in cancer molecular subtype classification. C_LIO_LIThe study highlights biases in specific omics modalities: SNP, RPPA, and Array exhibit higher variability and precision-recall imbalances, while RNASeq, miRNA, Exon, and CNV deliver more consistent and reliable results. C_LIO_LIML models (e.g., SVM, XGB, HGB) demonstrate strong performance on smaller datasets with fewer features, whereas DL models (e.g., ResNet18, CNN, NN, MLP, and DEEPGENE transformer) excel in handling high-dimensional datasets with large sample sizes. C_LIO_LIThe findings provide critical insights for developing robust, standardized AI pipelines for precision oncology, enhancing reproducibility and enabling meaningful cross-method comparisons. C_LI

bioinformatics↗