Search bioRxiv⌕ Search

Biology subjects

Kucukbenli, E.

Publications and source records attributed to Kucukbenli, E..

3 recordsLinked to original sources

Blind Virtual Screening at Scale: A Scalable End-to-End Pipeline for Blind Docking and Affinity Prediction

Accurate and scalable prediction of protein-ligand interactions remains a central challenge in computational drug discovery, especially when the binding site is unknown (i.e., blind docking). We present a high-throughput, end-to-end algorithm for virtual screening that combines DiffDock, a diffusion-based generative model for blind docking, with UniDock Vina, an algorithm for rapid scoring. We benchmarked this approach on the CASF-2016 and DUD-E datasets, analyzing pose quality, scoring accuracy, and screening performance. We find that competitive screening power can be achieved when generating and scoring as few as three poses and without pose refinement, which facilitates scalability. Notably, our method achieves 86.78% and 82.00% for the percent of actives among the top 1% and 10% of ranked ligands, respectively, when generating as few as three poses per protein-ligand pair. The workflow is scalable, supporting blind docking and affinity prediction at a mean throughput of 0.76 seconds per protein-ligand pair when generating 40 ligand poses in batched mode parallelized to 8 NVIDIA A100 80G GPUs. These results demonstrate that accurate, large-scale blind virtual screening is feasible and offers a practical solution for screening against novel or less characterized protein targets. Code is available at: https://github.com/xinyu-dev/blind-screening-benchmark

bioinformatics↗

PINDER: The protein interaction dataset and evaluation resource

Protein-protein interactions (PPIs) are fundamental to understanding biological processes and play a key role in therapeutic advancements. As deep-learning docking methods for PPIs gain traction, benchmarking protocols and datasets tailored for effective training and evaluation of their generalization capabilities and performance across real-world scenarios become imperative. Aiming to overcome limitations of existing approaches, we introduce PINDER, a comprehensive annotated dataset that uses structural clustering to derive non-redundant interface-based data splits and includes holo (bound), apo (unbound), and computationally predicted structures. PINDER consists of 2,319,564 dimeric PPI systems (and up to 25 million augmented PPIs) and 1,955 high-quality test PPIs with interface data leakage removed. Additionally, PINDER provides a test subset with 180 dimers for comparison to AlphaFold-Multimer without any interface leakage with respect to its training set. Unsurprisingly, the PINDER benchmark reveals that the performance of existing docking models is highly overestimated when evaluated on leaky test sets. Most importantly, by retraining DiffDock-PP on PINDER interface-clustered splits, we show that interface cluster-based sampling of the training split, along with the diverse and less leaky validation split, leads to strong generalization improvements.

bioinformatics↗

PLINDER: The protein-ligand interactions dataset and evaluation resource

Protein-ligand interactions (PLI) are foundational to small molecule drug design. With computational methods striving towards experimental accuracy, there is a critical demand for a well-curated and diverse PLI dataset. Existing datasets are often limited in size and diversity, and commonly used evaluation sets suffer from training information leakage, hindering the realistic assessment of method generalization capabilities. To address these shortcomings, we present PLIN-DER, the largest and most annotated dataset to date, comprising 449,383 PLI systems, each with over 500 annotations, similarity metrics at protein, pocket, interaction and ligand levels, and paired unbound (apo) and predicted structures. We propose an approach to generate training and evaluation splits that minimizes task-specific leakage and maximizes test set quality, and compare the resulting performance of DiffDock when retrained with different kinds of splits.

biochemistry↗