Search bioRxiv⌕ Search

Biology subjects

Mowrey, W. R.

Publications and source records attributed to Mowrey, W. R..

2 recordsLinked to original sources

Representation Learning of Human Disease Mechanisms for a Foundation Model in Rare and Common Diseases

A fundamental challenge in translational medicine is the computational modeling of complex human diseases to accelerate therapeutic development. Representation learning provides a powerful framework to address this, yet creating models that capture deep biological mechanisms remains a critical need. To this end, we propose a novel strategy that partitions the disease landscape into rare and non-rare categories, enabling systematic knowledge repurposing both within and between these groups. Here, we introduce Dis2Vec (Disease to Vector), a representation learning framework designed to operationalize this concept. Dis2Vec generates biologically grounded disease embeddings by learning from human genetic and phenotypic data, forming the foundation for Disease-Disease Association Learning (DDAL) and unsupervised disease clustering. We evaluate Dis2Vec representations in two downstream applications. First, we assess DDAL performance on a transfer learning benchmark designed to predict therapeutic transferability, using real-world drug repurposing investment decisions made in clinical trials. Second, unsupervised clustering analyses reveal shared biological mechanisms across diseases. By modeling the disease landscape in this way, Dis2Vec enhances translational research efficiency across both rare and non-rare diseases, advancing the development of foundational models for therapeutic science. Furthermore, Dis2Vec establishes a biologically grounded disease-representation and benchmarking layer that paves the way for trustworthy agentic biomedical AI systems in rare-disease indication expansion.

systems biology↗

Topology-Driven Negative Sampling Enhances Generalizability in Protein-Protein Interaction Prediction

Unraveling the human interactome to uncover disease-specific patterns and discover drug targets hinges on accurate protein-protein interaction (PPI) predictions. However, challenges persist in machine learning (ML) models due to a scarcity of quality hard negative samples, shortcut learning, and limited generalizability to novel proteins. Here, we introduce a novel approach for strategic sampling of protein-protein non-interactions (PPNIs) by leveraging higher-order network characteristics that capture the inherent complementarity-driven mechanisms of PPIs. Next, we introduce UPNA-PPI (Unsupervised Pre-training of Node Attributes tuned for PPI), a high throughput sequence-to-function ML pipeline, integrating unsupervised pretraining in protein representation learning with topological PPNI samples, capable of efficiently screening billions of interactions. UPNA-PPI improves PPI prediction generalizability and interpretability, particularly in identifying potential binding sites locations on amino acid sequences, strengthening the prioritization of screening assays and facilitating the transferability of ML predictions across protein families and homodimers. UPNA-PPI establishes the foundation for a fundamental negative sampling methodology in graph machine learning by integrating insights from network topology.

bioinformatics↗