Search bioRxiv⌕ Search

Biology subjects

Tam, H. H.

Publications and source records attributed to Tam, H. H..

2 recordsLinked to original sources

Blind Virtual Screening at Scale: A Scalable End-to-End Pipeline for Blind Docking and Affinity Prediction

Accurate and scalable prediction of protein-ligand interactions remains a central challenge in computational drug discovery, especially when the binding site is unknown (i.e., blind docking). We present a high-throughput, end-to-end algorithm for virtual screening that combines DiffDock, a diffusion-based generative model for blind docking, with UniDock Vina, an algorithm for rapid scoring. We benchmarked this approach on the CASF-2016 and DUD-E datasets, analyzing pose quality, scoring accuracy, and screening performance. We find that competitive screening power can be achieved when generating and scoring as few as three poses and without pose refinement, which facilitates scalability. Notably, our method achieves 86.78% and 82.00% for the percent of actives among the top 1% and 10% of ranked ligands, respectively, when generating as few as three poses per protein-ligand pair. The workflow is scalable, supporting blind docking and affinity prediction at a mean throughput of 0.76 seconds per protein-ligand pair when generating 40 ligand poses in batched mode parallelized to 8 NVIDIA A100 80G GPUs. These results demonstrate that accurate, large-scale blind virtual screening is feasible and offers a practical solution for screening against novel or less characterized protein targets. Code is available at: https://github.com/xinyu-dev/blind-screening-benchmark

bioinformatics↗

ComboPath: An ML system for predicting drug combination effects with superior model specification

Drug combinations have been shown to be an effective strategy for cancer therapy, but identifying beneficial combinations through experiments is labor-intensive and expensive [Mokhtari et al., 2017]. Machine learning (ML) systems that can propose novel and effective drug combinations have the potential to dramatically improve the efficiency of combinatoric drug design. However, the biophysical parameters of drug combinations are degenerate, making it difficult to identify the ground truth of drug interactions even given experimental data of the highest quality available. Existing ML models are highly underspecified to meet this challenge, leaving them vulnerable to producing parameters that are not biophysically realistic and harming generalization. We have developed a new ML model, "ComboPath", aimed at a novel ML task: to predict interpretable cellular dose response surface of a two-drug combination based on each drugs interactions with their known protein targets. ComboPath incorporates a biophysically-motivated intermediate parameterization with prior information used to improve model specification. This is the first ML model to nominate beneficial drug combinations while simultaneously reconstructing the dose response surface, providing insight on both the potential of a drug combination and its optimal dosing for therapeutic development. We show that our models were able to accurately reconstruct 2D dose response surfaces across held out combination samples from the largest available combinatoric screening dataset while substantially improving model specification for key biophysical parameters.

biophysics↗