Search bioRxiv⌕ Search

Biology subjects

Pulido, I.

Publications and source records attributed to Pulido, I..

4 recordsLinked to original sources

How many crystal structures do you need to trust your docking results?

Structure-based drug discovery relies on the prediction of protein-bound poses of new molecule designs, the accuracy of which can impact downstream prioritization. While it is expected that crystal structures of similar molecules would provide the best template for predicting the poses of new designs, the time and cost required motivates identifying a point of diminishing returns for collecting new structures. Using 403 crystal structures of SARS-CoV-2 main protease from the open science COVID Moonshot project, we explore the tradeoff between the cost and utility of obtaining crystal structures for accurately predicting poses of designed molecules. We observe that similar reference ligands enable superior pose prediction and show that success plateaus after approximately five crystal structures per generic Bemis-Murcko scaffold, exceeding 95% for the campaign's lead series. This work provides practical recommendations for resource allocation in structure-enabled drug discovery campaigns.

biophysics↗

A Structure-Based Computational Pipeline for Broad-Spectrum Antiviral Discovery

The rapid emergence of viruses with pandemic potential continues to pose a threat to public health worldwide. With the typical drug discovery pipeline taking an average of 5-10 years to reach clinical readiness, there is an urgent need for strategies to develop broad-spectrum antivirals that can target multiple viral family members and variants of concern. We present a structure-based computational pipeline designed to identify and evaluate broad-spectrum inhibitors across viral family members for a given target in order to support spectrum breadth assessment and prioritization in lead optimization programs. This pipeline comprises three key steps: (1) an automated search to identify viral sequences related to a specified target construct, (2) pose prediction leveraging any available structural data, and (3) scoring of protein-ligand complexes to estimate antiviral activity breadth. The pipeline is implemented using the drugforge package: an open-source toolkit for structure-based antiviral discovery. To validate this framework, we retrospectively evaluated two overlapping datasets of ligands bound to the SARS-CoV-2 and MERS-CoV main protease (Mpro), observing useful predictive power with respect to experimental binding affinities. Additionally, we screened known SARS-CoV-2 Mpro inhibitors against a panel of human and non-human coronaviruses, demonstrating the potential of this approach to assess broad-spectrum antiviral activity. Our computational strategy aims to accelerate the identification of antiviral therapies for current and emerging viruses with pandemic potential, contributing to global preparedness for future outbreaks.

bioinformatics↗

Lessons learned during the journey of data: from experiment to model for predicting kinase affinity, selectivity, polypharmacology, and resistance

Recent advances in machine learning (ML) are reshaping drug discovery. Structure-based ML methods use physically-inspired models to predict binding affinities from protein:ligand complexes. These methods promise to enable the integration of data for many related targets, which addresses issues related to data scarcity for single targets and could enable generalizable predictions for a broad range of targets, including mutants. In this work, we report our experiences in building KinoML, a novel framework for ML in target-based small molecule drug discovery with an emphasis on structure-enabled methods. KinoML focuses currently on kinases as the relative structural conservation of this protein superfamily, particularly in the kinase domain, means it is possible to leverage data from the entire superfamily to make structure-informed predictions about binding affinities, selectivities, and drug resistance. Some key lessons learned in building KinoML include: the importance of reproducible data collection and deposition, the harmonization of molecular data and featurization, and the choice of the right data format to ensure reusability and reproducibility of ML models. As a result, KinoML allows users to easily achieve three tasks: accessing and curating molecular data; featurizing this data with representations suitable for ML applications; and running reproducible ML experiments that require access to ligand, protein, and assay information to predict ligand affinity. Despite KinoML focusing on kinases, this framework can be applied to other proteins. The lessons reported here can help guide the development of platforms for structure-enabled ML in other areas of drug discovery.

biophysics↗

Identifying and overcoming the sampling challenges in relative binding free energy calculations of a model protein:protein complex

Relative alchemical binding free energy calculations are routinely used in drug discovery projects to optimize the affinity of small molecules for their drug targets. Alchemical methods can also be used to estimate the impact of amino acid mutations on protein:protein binding affinities, but these calculations can involve sampling challenges due to the complex networks of protein and water interactions frequently present in protein:protein interfaces. We investigate these challenges by extending a GPU-accelerated open-source relative free energy calculation package (Perses) to predict the impact of amino acid mutations on protein:protein binding. Using the well-characterized model system barnase:barstar, we describe analyses for identifying and characterizing sampling problems in protein:protein relative free energy calculations. We find that mutations with sampling problems often involve charge-changes, and inadequate sampling can be attributed to slow degrees of freedom that are mutation-specific. We also explore the accuracy and efficiency of current state-of-the-art approaches--alchemical replica exchange and alchemical replica exchange with solute tempering--for overcoming relevant sampling problems. By employing sufficiently long simulations, we achieve accurate predictions (RMSE 1.61, 95% CI: [1.12, 2.11] kcal/mol), with 86% of estimates within 1 kcal/mol of the experimentally-determined relative binding free energies and 100% of predictions correctly classifying the sign of the changes in binding free energies. Ultimately, we provide a model workflow for applying protein mutation free energy calculations to protein:protein complexes, and importantly, catalog the sampling challenges associated with these types of alchemical transformations. Our free open-source package (Perses) is based on OpenMM and available at https://github.com/choderalab/perses.

biophysics↗