Search bioRxiv⌕ Search

Biology subjects

Tashchilova, A.

Publications and source records attributed to Tashchilova, A..

1 recordsLinked to original sources

DrugForm-DTA: Towards real-world drug-target binding Affinity Model

Drug-target affinity (DTA) prediction is a fundamental problem in drug discovery. Computational methods for predicting DTA can greatly assist drug design by decreasing the search space and reducing the number of protein-ligand complexes with low affinity. Currently DTA approaches often do not require protein three-dimensional (3D) structural information, which is often not accessible. In this study we present the DrugForm-DTA model, which uses only structure-less representations of ligand and protein. It is a Transformer-based neural network with protein encoding based on ESM, and small molecule ligand encoding obtained with Chemformer. We evaluated the model on standard benchmarks Davis and KIBA, and revealed superior performance of DrugForm-DTA with best result for KIBA (MSE=0.117). Moreover, we developed a ready-to-use model using BindingDB dataset that was subjected to high-quality filtering and transformation. Overall, our method predicts drug-target affinity values with a confidence level comparable to a single in-vitro experiment. Also, we compared DrugForm-DTA against molecular modeling methods and revealed higher efficacy of the developed model for drug-target affinity predictions. Our investigation provides a high accuracy neural network model with performance comparable to experi-mental measurements, filtered and reassessed BindingDB dataset for further usage, and demonstrates outstanding applicability of the proposed method for DTA prediction. Author summaryPredicting drug-target binding affinity is a crucial task in computational drug discovery. Here we present DrugForm-DTA, a ready-to-use Transformer-based model which requires only protein amino acid sequence and ligand SMILES as inputs. The model shows excellent results at commonly used benchmarks. We believe that the success in solving machine learning tasks depends on data quality and quantity, so the main focus of our work is the training data, but not a sophisticated neural architecture. We used BindingDB database as the source of experimental affinity measurements and prepared a refined and purified dataset, suitable for training machine learning models. Trained on this dataset, the DrugForm-DTA model demonstrates accuracy, comparable to a single in-vitro experiment. We also reveal that our model outperforms molecular modeling methods in estimating binding affinity. Both the prepared dataset and the trained model is published and freely available.

bioinformatics↗