Interrogating contrastive learning embeddings for structure-based virtual screening: a case study on DrugCLIP
Virtual screening has become central to early-stage drug discovery, and structure-based approaches have recently been reframed as a retrieval problem through contrastive learning methods such as DrugCLIP, which project protein pockets and ligands into a shared embedding space. However, what these abstract representations exactly encode, and how they relate to conventional notions of structural and chemical similarity, remains unclear. Here we present a systematic dissection of DrugCLIP's latent space. We show that its pocket embeddings, despite not being explicitly trained for the task, set a new state of the art in pocket similarity search while running over 100 times faster than existing structural descriptors, and that this embedding space is structurally coherent and robust to conformational variation. Ligand embeddings, by contrast, encode a pocket-aware notion of chemical similarity that only partially mirrors fingerprint-based measures. Using a rigorous de-leakage benchmark, we further show that DrugCLIP generalises to unseen proteins and chemistries rather than memorising training data, recovering the correct bound ligand within the top 1% of 50,000 candidates for 55-75% of novel targets. Performance nonetheless declines under increasingly realistic screening conditions, a drop attributable to sidechain reorientation across apo, holo and AlphaFold-derived structures, and to residue mismatch when using predicted pockets. These findings clarify the practical boundaries of DrugCLIP's applicability, identify pocket prediction accuracy as a key factor for improving performance, and offer a transferable framework for interpreting the latent spaces of related contrastive pocket-ligand encoders. Together, these results support the improvement of existing methods and the development of a new generation of contrastive screening approaches.