Search bioRxiv⌕ Search

Biology subjects

Lambden, S.

Publications and source records attributed to Lambden, S..

2 recordsLinked to original sources

Clinical Trial and Ontology-Derived Positive and Negative Benchmark Datasets for Drug Repurposing Across Rare Diseases

Evaluating the potential applications of a medicine is a fundamental challenge in drug development. There is a lack of standardized, decision-oriented benchmarks that test whether computational models can generalize therapeutic hypotheses across diseases in ways that reflect real-world pharmaceutical investment decision making. To address this gap, we introduce two complementary resources: the Indication Expansion Investment Decision Network (IxIDN) and the Orphanet Rare Disease Ontology Negative-network (ORDON). IxIDN is a clinical-trial-derived positive benchmark constructed by projecting drug-disease associations from pharmaceutical clinical trials into a disease-disease network; each edge connects disease pairs that have entered clinical trials for the same drug, thereby capturing cases when concrete indication-expansion decisions have been made. The current release contains 574 rare diseases and 5,336 edges. In contrast, ORDON serves as a stringent, biology-aware negative benchmark derived from the authoritative Orphanet Rare Disease Ontology. It identifies maximally distant disease pairs according to curated hierarchical structure and genetics-linked inheritance patterns, providing 793 rare diseases and 5,000 edges that represent high-separation negative candidates across therapeutic areas. Together, IxIDN and ORDON enable rigorous cross-evidence generalization from clinical trials to disease ontology, testing for Disease- Disease Association Learning (DDAL), a core task for mechanism-centered drug repurposing and indication expansion. All data are publicly available with detailed metadata, enabling reproducible evaluation of models on transparent, decision-relevant benchmarks.

Systems Biology↗

Representation Learning of Human Disease Mechanisms for a Foundation Model in Rare and Common Diseases

A fundamental challenge in translational medicine is the computational modeling of complex human diseases to accelerate therapeutic development. Representation learning provides a powerful framework to address this, yet creating models that capture deep biological mechanisms remains a critical need. To this end, we propose a novel strategy that partitions the disease landscape into rare and non-rare categories, enabling systematic knowledge repurposing both within and between these groups. Here, we introduce Dis2Vec (Disease to Vector), a representation learning framework designed to operationalize this concept. Dis2Vec generates biologically grounded disease embeddings by learning from human genetic and phenotypic data, forming the foundation for Disease-Disease Association Learning (DDAL) and unsupervised disease clustering. We evaluate Dis2Vec representations in two downstream applications. First, we assess DDAL performance on a transfer learning benchmark designed to predict therapeutic transferability, using real-world drug repurposing investment decisions made in clinical trials. Second, unsupervised clustering analyses reveal shared biological mechanisms across diseases. By modeling the disease landscape in this way, Dis2Vec enhances translational research efficiency across both rare and non-rare diseases, advancing the development of foundational models for therapeutic science. Furthermore, Dis2Vec establishes a biologically grounded disease-representation and benchmarking layer that paves the way for trustworthy agentic biomedical AI systems in rare-disease indication expansion.

systems biology↗