Search bioRxiv⌕ Search

Biology subjects

Martinez Leon, A.

Publications and source records attributed to Martinez Leon, A..

2 recordsLinked to original sources

BindFlow: a free, user-friendly pipeline for absolute binding free energy calculations using free energy perturbation or MM(PB/GB)SA

We present BindFlow, a Python-based software for automated absolute binding free energy (ABFE) calculations at the free energy perturbation (FEP) or at the molecular mechanics Poisson-Boltzmann/generalized Born surface area [MM(PB/GB)SA] level of theory. BindFlow is free, open-source, user-friendly, easily customizable, runs on work-stations or distributed computing platforms, and provides extensive documentation and tutorials. BindFlow uses GROMACS as molecular dynamics engine and provides built-in support for the small-molecule force fields GAFF, OpenFF, and Espaloma. We test BindFlow by computing affinities for 139 ligand/target pairs, involving eight different targets including six soluble proteins, one membrane protein and one non-protein host- guest system. Quantified by Pearson, Kendall, and Spearman correlations coefficients, we find that the agreement of BindFlow predictions with experiments are overall similar to gold standards in the field. Interestingly, we find that MM(PB/GB)SA achieves correlations that, for some systems and force fields, approach those obtained with FEP, while requiring only a fraction of the computational cost. This study establishes BindFlow as a validated and accessible tool for ABFE calculations.

biophysics↗

Accelerating ligand discovery by combining Bayesian optimization with MMGBSA-based binding affinity calculations

Predicting protein-ligand binding affinity with high accuracy is critical in structure-based drug discovery. While docking methods offer computational efficiency, they often lack the precision required for reliable affinity ranking. In contrast, molecular dynamics (MD)-based approaches such as MMGBSA provide more accurate binding free energy estimates but are computationally intensive, limiting their scalability. To address this trade-off, we introduce an active learning framework that automates molecule selection for docking and MD simulations, replacing manual expert-driven decisions with a data-efficient, model-guided strategy. Our approach integrates fixed -- partly pre-trained deep learning -- molecular embeddings (MolFormer, ChemBERTa-2, and Morgan fingerprints) with adaptive regression models (e.g. Bayesian Ridge and Random Forest) to iteratively improve binding affinity predictions. We evaluate this approach retro-spectively on a new dataset of 59,356 chemically diverse compounds from ZINC-22 targeting the MCL1 protein using both AutoDock Vina and MMGBSA binding free energy scores. Our results show that incorporating MMGBSA scores into the active learning loop significantly enhances performance, recovering 79.9% of the top 1% binders in the whole dataset, compared to only 6.7% when using docking scores alone. Notably, MMGBSA exhibits a stronger correlation with experimental binding affinities than AutoDock Vina on our dataset and enables more accurate ranking of candidate compounds in a runtime efficient way. Furthermore, we demonstrate that a one-at-a-time acquisition active learning strategy consistently outperforms traditional batched acquisition, the latter achieving just 78.4% recovery with MolFormer and Bayesian Ridge. These findings underscore the potential of integrating deep learning-based molecular representations with MD-level accuracy in an active learning framework, offering a scalable and efficient path to accelerate virtual screening and improve hit identification in drug discovery.

bioinformatics↗