Search bioRxiv⌕ Search

Biology subjects

Rayakar, A. A.

Publications and source records attributed to Rayakar, A. A..

2 recordsLinked to original sources

Integrating Multi-Structure Covalent Docking with Machine Learning Consensus Scoring Enhances Virtual Screening of Human Acetylcholinesterase Inhibitors

Acetylcholinesterase (AChE) inhibition is a key mechanism in the treatment of neurodegenerative diseases and in counteracting toxic exposures to pesticides and nerve agents. However, virtual screening of AChE remains challenging due to the enzymes structural flexibility and the chemical diversity of its covalently binding inhibitors. In this study, we developed an in silico protocol that integrates multi-structure covalent docking and machine learning (ML) consensus scoring to improve the prediction of AChE inhibitors. We analyzed 65 ligand-bound (holo) human AChE crystal structures using hierarchical clustering to identify four representative conformations, along with one high-resolution apo structure, for multi-structure docking. A curated library of 412 organophosphate and carbamate inhibitors was then docked covalently and non-covalently into each receptor conformation. The resulting docking scores were evaluated against inhibitors experimental logIC50 values using Spearmans rank correlation coefficient (r). Covalent docking outperformed non-covalent docking (r values up to 0.54 vs 0.18), and our ML consensus model trained on the five structures covalent docking scores achieved the highest predictive accuracy (r = 0.70), surpassing all single-structure and conventional consensus baselines. Chemical cluster analysis revealed structure-activity trends based on ligand flexibility, polarity, and aromaticity. SHapley Additive exPlanations analysis highlighted the ML consensus models ability to flexibly distribute the influence each structures scores played on its predictions. It identified and exploited relationships based on its training dataset that would be difficult to anticipate through a manual analysis of individual structures docking performance metrics. This framework is broadly applicable to other covalently targeted proteins, offering a generalizable and interpretable strategy for data-driven covalent inhibitor discovery.

bioinformatics↗

Deep contrastive feature compression with classical machine learning enables ligand discovery through efficient triage of large chemical libraries

Improving in silico compound-protein interaction (CPI) predictability is critical for productive drug discovery. Current deep learning approaches largely rely on end-to-end models trained on limited labeled CPI datasets, overlooking the representational power of large-scale biochemical foundation models. We present COMRADE (Contrastive Multirepresentation Accelerated Docking Engine), a hybrid virtual screening framework that accelerates docking by triaging compounds using CE-Screen (Contrastive Embedding-Screen). CE-Screen leverages seven high-dimensional pretrained representations - including those from protein language models and molecular transformers, along with an original physics-based interaction potential encoding - for rapid first-pass screening [~]100x faster than docking. Its contrastive compression neural network maps these inputs onto a single compact, discriminative representation optimized for CPI prediction via a lightweight ensemble classifier. CE-Screen outperforms state-of-the-art end-to-end models by up to 111.11% on retrospective benchmarks and is successfully used to triage [~]10.8 million compounds against five targets, yielding novel hits for each one - including a new scaffold for the branched-chain ketoacid dehydrogenase kinase (BCKDK), an understudied yet high-value target in metabolic disease and oncology.

bioinformatics↗