Seq2Pocket: Augmenting protein language models for spatially consistent binding site prediction
Protein-ligand binding site prediction (LBS) is important for many areas of structural biology and molecular modeling, where, as in other tasks, protein language models (pLMs) have shown a great promise. In their application to LBS, the pLM classifies each amino acid as binding or not, but translating these predictions into three-dimensional binding pockets remains challenging; in particular, residue-centric predictions tend to produce spatially fragmented pockets. We present Seq2Pocket, a methodology for pocket-level LBS prediction that combines pLM finetuning, data enhancement, and structure-aware post-processing. First, we introduce sc-PDBenhanced, an extended training dataset that augments sc-PDB with additional small-molecule and ion-binding sites, improving coverage of non-obvious and small pockets. Second, we employ an embedding-supported smoothing classifier to refine residue-level predictions. Third, we define the Pocket Fragmentation Index and use it to select a clustering approach that preserves a consistent mapping between predictions and ground-truth pockets. We evaluate Seq2Pocket on two tasks: general binding site prediction using the LIGYSIS benchmark and cryptic binding site prediction using CryptoBench benchmark. Across both benchmarks, the proposed methodology achieves state-of-the-art performance. In particular, for general binding site prediction on the LIGYSIS benchmark, it improves distance-center-to-center recall by up to 12%, outperforming existing predictors. We believe that these findings contribute to more reliable evaluation practices in ligand binding site prediction, highlight the importance of training data curation, and provide pocket-level prediction tool that are better suited for downstream applications such as drug discovery.