Search bioRxiv⌕ Search

Biology subjects

Van Damme, M.

Publications and source records attributed to Van Damme, M..

3 recordsLinked to original sources

On improving experimental binding affinity predictions with synthetic data

The success of deep learning binding affinity prediction models depends critically on expanding experimental data with reliable synthetic data. We extend the Structurally Augmented IC50 Repository (SAIR) with {approx}80K absolute free energy perturbation (AFEP) calculations and present two distinct data splits, SAIR-FEP and SAIR-OOD (out-of-distribution), to simulate realistic drug discovery scenarios. We compare sequence-based proteochemometric (PCM) models and state-of-the-art, structure-based deep learning models and demonstrate that PCM models can be enhanced by physics-based descriptors. While structure-based deep learning methods capture finer geometric detail, their performance is highly sensitive to the input structure. By filtering for high-confidence, co-folded complexes, we show that the performance improves predictably, whereas training on all complexes blindly does not yield performance gains. Finally, using the SAIR-OOD split, we demonstrate that simultaneous training on synthetic and experimental data improves performance on publicly available, experimental benchmarks. These results provide a clear strategy for using synthetic data to advance experimental binding affinity predictions.

molecular biology↗

Longitudinal dynamics of organ-specific proteomic aging clocks over a decade of midlife

Organ-specific proteomic clocks are promising tools for quantifying heterogeneity in biological aging, but their longitudinal behavior remains largely unexplored. Here, we analyzed paired plasma proteomic profiles with 10-year follow-up in middle-aged adults (n= 1,250) to evaluate their longitudinal properties. Cross-sectional associations of protein concentrations with age mirrored average longitudinal trajectories, validating the common cross-sectional training of clocks. Organ-specific age acceleration was moderately stable over the decade, and aging across organs progressed in parallel, with the immune and adipose systems acting as central hubs and early cardiorespiratory aging predicting downstream metabolic aging. Critically, longitudinal changes in predicted age tracked subclinical risk factor alterations. In women, the menopausal transition dominated the aging landscape and was associated with multi-organ age acceleration. Medication initiation altered clocks through specific drug-targeted proteins (such as renin and APOB) rather than generalized organ aging. Together, these findings position organ-specific proteomic clocks as interpretable, dynamic indicators of aging and organ health.

systems biology↗

SAIR: Enabling deep learning for protein-ligand lnteractions with a synthetic structural dataset

AO_SCPLOWBSTRACTC_SCPLOWAccurate prediction of protein-ligand binding affinities remains a cornerstone problem in drug discovery. While binding affinity is inherently dictated by the 3D structure and dynamics of protein-ligand complexes, current deep learning approaches are limited by the lack of high-quality experimental structures with annotated binding affinities. To address this limitation, we introduce the Struc-turally Augmented IC50 Repository (SAIR), the largest publicly available dataset of protein-ligand 3D structures with associated activity data. The dataset com-prises 5, 244, 285 structures across 1, 048, 857 unique protein-ligand systems, cu-rated from the ChEMBL and BindingDB databases, which were then computa-tionally folded using the Boltz-1x model. We provide a comprehensive charac-terization of the dataset, including distributional statistics of proteins and ligands, and evaluate the structural fidelity of the folded complexes using PoseBusters. Our analysis reveals that approximately 3% of structures exhibit physical anoma-lies, predominantly related to internal energy violations. As an initial demon-stration, we benchmark several binding affinity prediction methods, including empirical scoring functions (Vina, Vinardo), a 3D convolutional neural network (Onionnet-2), and a graph neural network (AEV-PLIG). While machine learning-based models consistently outperform traditional scoring function methods, nei-ther exhibit a high correlation with ground truth affinities, highlighting the need for models specifically fine-tuned to synthetic structure distributions. This work provides a foundation for developing and evaluating next-generation structure and binding-affinity prediction models and offers insights into the structural and phys-ical underpinnings of protein-ligand interactions. The dataset can be found at https://www.sandboxaq.com/sair.

biochemistry↗