Search bioRxiv⌕ Search

Biology subjects

Prol-Castelo, G.

Publications and source records attributed to Prol-Castelo, G..

3 recordsLinked to original sources

Interpretable Forecasting of Kidney Cancer Progression via Generative AI and Symbolic Reasoning

Predicting cancer stage progression from omics data, and deriving molecular insight into the mechanisms driving it, remains a major challenge, owing in part to the lack of adequate longitudinal data and the interpretability limitations of current forecasting models. Large cancer datasets such as TCGA capture patient profiles cross-sectionally rather than longitudinally, complicating timely treatment decisions as tumors become more invasive. Deep neural networks typically used for forecasting, such as LSTMs, compound this problem by remaining largely opaque and offering clinicians no straightforward way to audit their predictions. Clear cell renal cell carcinoma (ccRCC) illustrates the clinical stakes of both challenges. Five-year survival falls from over 94% at stage I to 28% at stage IV, yet early-stage tumors are often managed under active surveillance, a strategy constrained by sparse molecular evidence of progression risk. Detecting progression in time, meanwhile, demands forecasts clinicians can interpret and trust, not black-box predictions. We address both challenges by combining generative and symbolic AI: a Variational Autoencoder trained on bulk RNA-Seq profiles of 530 TCGA ccRCC patients generates synthetic pseudo-time trajectories that overcome the absence of longitudinal data, while a symbolic rule-induction framework (ASAL) learns finite-state automata from these trajectories, encoding stage transition as human-readable Boolean conditions over gene expression, which a complex event forecasting system (Wayeb) converts into probabilistic forecasts of stage advancement. An independent XGBoost classifier trained on real patients (F1 score = 0.71-0.81) shows a gradual early-to-late probability shift along the synthetic trajectories, absent in non-progressing control trajectories. Pathway enrichment of those trajectories reveals stage-dependent changes in established kidney cancer-related processes, including the TCA cycle and DNA repair. Finally, our symbolic forecaster nearly matches an LSTM baseline (macro F1 = 0.928 vs. 0.964), while additionally offering an inspectable rule set and a probability distribution over transition timing rather than a single opaque score. This work shows that generative and symbolic AI, paired together, can turn cross-sectional cohorts into a transparent, forecast-oriented framework for modeling disease progression, demonstrated here in ccRCC.

bioinformatics↗

10 Years of Variational Autoencoder: Insights from Cancer Temporal Progression Studies, a Systematic Literature Review

Deep learning methods, including deep representation learning (DRL) approaches such as variational au-toencoders (VAEs), have been widely applied to cancer omics data to address the high dimensionality of these datasets. Despite remarkable advances, cancer remains a complex and dynamic disease that is challenging to study, and the temporal resolution of cancer progression captured by omics-based studies remains limited. In this systematic literature review, we explore the use of DRL, particularly the VAE, in cancer omics studies for modeling time-related processes, such as tumor progression and evolutionary dynamics. Our work reveals that these methods most commonly support subtyping, diagnosis, and prognosis in this context, but rarely emphasize temporal information. We observed that the scarcity of longitudinal omics data currently limits deeper temporal analyses that could enhance these applications. We propose that applying the VAE as a generative model to study cancer in time, for example, focusing on cancer staging, could lead to meaningful advancements in our understanding of the disease. Biographical NoteO_LIGuillermo Prol-Castelo is a PhD student at the Barcelona Supercomputing Center and Universitat Pompeu Fabra, where he works on the application of deep learning methods to cancer studies. C_LIO_LIDavide Cirillo is the head of the Machine Learning for Biomedical Research Unit at the Barcelona Supercomputing Center. He is an expert in predictive modeling for Precision Medicine using Network Biology and Machine Learning. C_LIO_LIAlfonso Valencia is the principal investigator of the Computational Biology Group at the Barcelona Supercomputing Center. He is a leading expert in protein coevolution, disease networks and modelling cellular systems. C_LIO_LIThe Barcelona Supercomputing Center is a public research center that provides high-performance computing infrastructure to support scientific research in a wide range of fields, including life sciences. C_LI Key PointsO_LIThere is a growing interest on the application of deep learning methods, such as Deep Representation Learning (DRL), to cancer studies. C_LIO_LICancer is a complex and dynamic disease, whose temporal dynamics are not yet fully captured in omics-based studies. C_LIO_LImong DRL methods, the Variational Autoencoder (VAE) using omics-based data has been widely used in cancer studies, particularly for subtyping, diagnosis, and prognosis. C_LIO_LIThe temporal aspects of cancer progression are often insufficiently captured in omics-based studies, primarily due to the scarcity of longitudinal data. C_LIO_LIApplying the VAE as a generative model to study cancer in time, such as focusing on cancer staging, could lead to significant advancements in our understanding of cancer. C_LI

bioinformatics↗

Exploring the Boundaries of Medulloblastoma Subgroups with Synthetic Data Generation

Medulloblastoma is a childhood brain tumor traditionally classified into four molecular subgroups. Recent evidence suggests that Groups 3 and 4 represent a biological continuum rather than distinct entities, a paradigm shift with significant implications for understanding disease biology and treatment strategies. Nevertheless, assessing this hypothesis is challenging mainly due to data scarcity. In this study, we analyze the largest available transcriptomics dataset to provide compelling evidence for the existence of an intermediate subgroup between Groups 3 and 4, characterized by distinct molecular features. To overcome limitations posed by data scarcity, we employ synthetic data generation using a Variational Autoencoder and apply explainability techniques to identify key relationships between gene expression and disease subgroups. Furthermore, by incorporating Machine Learning Fairness approaches, we demonstrate that overlooking this intermediate subgroup can result in treatment disparities. Our findings are further supported by both existing and newly proposed studies using diverse datasets and methodologies, including graphbased analyses and multi-scale simulations, underscoring the robustness and reproducibility of our results. This study demonstrates the potential of synthetic data generation to refine rare disease subtyping and advance our understanding of the underlying biological mechanisms. Keywords: Medulloblastoma, pediatric cancer, representation learning, autoencoder, synthetic data

bioinformatics↗