Search bioRxiv⌕ Search

Biology subjects

Bandi, S.

Publications and source records attributed to Bandi, S..

3 recordsLinked to original sources

On improving experimental binding affinity predictions with synthetic data

The success of deep learning binding affinity prediction models depends critically on expanding experimental data with reliable synthetic data. We extend the Structurally Augmented IC50 Repository (SAIR) with {approx}80K absolute free energy perturbation (AFEP) calculations and present two distinct data splits, SAIR-FEP and SAIR-OOD (out-of-distribution), to simulate realistic drug discovery scenarios. We compare sequence-based proteochemometric (PCM) models and state-of-the-art, structure-based deep learning models and demonstrate that PCM models can be enhanced by physics-based descriptors. While structure-based deep learning methods capture finer geometric detail, their performance is highly sensitive to the input structure. By filtering for high-confidence, co-folded complexes, we show that the performance improves predictably, whereas training on all complexes blindly does not yield performance gains. Finally, using the SAIR-OOD split, we demonstrate that simultaneous training on synthetic and experimental data improves performance on publicly available, experimental benchmarks. These results provide a clear strategy for using synthetic data to advance experimental binding affinity predictions.

molecular biology↗

SAIR: Enabling deep learning for protein-ligand lnteractions with a synthetic structural dataset

AO_SCPLOWBSTRACTC_SCPLOWAccurate prediction of protein-ligand binding affinities remains a cornerstone problem in drug discovery. While binding affinity is inherently dictated by the 3D structure and dynamics of protein-ligand complexes, current deep learning approaches are limited by the lack of high-quality experimental structures with annotated binding affinities. To address this limitation, we introduce the Struc-turally Augmented IC50 Repository (SAIR), the largest publicly available dataset of protein-ligand 3D structures with associated activity data. The dataset com-prises 5, 244, 285 structures across 1, 048, 857 unique protein-ligand systems, cu-rated from the ChEMBL and BindingDB databases, which were then computa-tionally folded using the Boltz-1x model. We provide a comprehensive charac-terization of the dataset, including distributional statistics of proteins and ligands, and evaluate the structural fidelity of the folded complexes using PoseBusters. Our analysis reveals that approximately 3% of structures exhibit physical anoma-lies, predominantly related to internal energy violations. As an initial demon-stration, we benchmark several binding affinity prediction methods, including empirical scoring functions (Vina, Vinardo), a 3D convolutional neural network (Onionnet-2), and a graph neural network (AEV-PLIG). While machine learning-based models consistently outperform traditional scoring function methods, nei-ther exhibit a high correlation with ground truth affinities, highlighting the need for models specifically fine-tuned to synthetic structure distributions. This work provides a foundation for developing and evaluating next-generation structure and binding-affinity prediction models and offers insights into the structural and phys-ical underpinnings of protein-ligand interactions. The dataset can be found at https://www.sandboxaq.com/sair.

biochemistry↗

Multi-Modal Large Language Model Enables All-Purpose Prediction of Drug Mechanisms and Properties

Accurately predicting the mechanisms and properties of potential drug molecules is essential for advancing drug discovery. However, traditional methods often require the development of specialized models for each specific prediction task, resulting in inefficiencies in both model training and integration into work-flows. Moreover, these approaches are typically limited to predicting pharmaceutical attributes represented as discrete categories, and struggle with predicting complex attributes that are best described in free-form texts. To address these challenges, we introduce DrugChat, a multi-modal large language model (LLM) designed to provide comprehensive predictions of molecule mechanisms and properties within a unified framework. DrugChat analyzes the structure of an input molecule along with users queries to generate comprehensive, free-form predictions on drug indications, pharmacodynamics, and mechanisms of action. Moreover, DrugChat supports multi-turn dialogues with users, facilitating interactive and in-depth exploration of the same molecule. Our extensive evaluation, including assessments by human experts, demonstrates that DrugChat significantly outperforms GPT-4 and other leading LLMs in generating accurate free-form predictions, and exceeds state-of-the-art specialized prediction models.

pharmacology and toxicology↗