Search bioRxiv⌕ Search

Biology subjects

Gingrich, P. W.

Publications and source records attributed to Gingrich, P. W..

2 recordsLinked to original sources

PocketBagger: Generalizable pocket druggability prediction via positive-unlabeled learning

Abstract SummaryReliable structure-based prediction of small-molecule druggability is hindered by a fundamental labeling problem. Experimentally confirmed liganded sites (positives) are observable, but credible "undruggable" pockets (negatives) are almost impossible to define. Standard supervised machine learning consequently relies on arbitrary definitions of undruggable, leading to bias and false negatives. Here we introduce PocketBagger, a positive-unlabeled (PU) learning framework for pocket druggability prediction trained exclusively on experimentally determined Protein Data Bank1 (PDB) structures. PocketBagger uses PU bagging to learn key features associated with reliable druggable pockets and considers all remaining pockets in the structurally characterized proteome as unlabeled. We demonstrate the capability of PocketBagger through the training of a simple Random Forest classifier and demonstrate its power in recall (0.804), even when challenged with increasingly difficult generalizability assessments and entire protein-family hold outs. We benchmark and demonstrate the added value of PU learning by comparing PocketBagger to a leading deep-learning predictor. However, PocketBagger is intended to be used as a framework for any model architecture. Along with the code, the data generated by PocketBagger are deployed in canSAR.ai, providing scalable, generalizable pocket druggability predictions to the drug discovery community.

bioinformatics↗

A Chemoproteomic Atlas of the Human Purine Interactome for Regioselective Ligand Discovery

Purines are essential bioactive molecules that interact with a large fraction of the human proteome. Despite their importance, the scope of actionable purine-binding pockets for ligand discovery remains limited. Here, we developed a quantitative chemoproteomics platform using sulfonyl-purine (SuPUR) chemistry to produce a massive and functional map of the human purine interactome. The SuPUR platform captured 31,000+ targetable tyrosine and lysine sites, representing the most comprehensive beyond cysteine chemoproteomics database for enabling protein ligand discovery. SuPUR ligands that bind through a regioselective fashion serve as enabling starting points for developing potent (nanomolar) and proteome-wide-selective modulators of enzymatic and protein-protein interaction function. Phenotypic screening identified a site-specific (Y237) and regioselective SuPUR ligand of ACAT2 to reveal an unexpected metabolic dependency in cancer cells. A crystal structure of SuPUR ligand-bound ACAT2 revealed the purine group binds deep in the CoA pocket forming key interactions with catalytic residues via a water bridge to guide future structure-based ligand design.

biochemistry↗