bioRxiv · 10.64898/2026.08.05.742817
Fast retrieval of structurally similar antibodies from large sequence databases with AbSLang
Abstract
The first steps in antibody therapeutic discovery involve identification of sequences with desirable binding properties. A way of finding these lead molecules is through the search of large sequence databases. Current methods, due to the size of databases, rely on germline or complementarity-determining-region (CDR) sequence identities, overlooking structurally similar antibodies with divergent sequences which can have identical binding properties . To address this, we introduce AbSLang, a model trained for pairwise CDR RMSD prediction using a contrastive learning approach. We demonstrate that AbSLang has comparable accuracy to exact RMSD calculation after explicit structure prediction with state-of-the-art models. Building on this model, we implemented AbSLang-search, a pipeline for retrieval of structurally similar antibodies from large sequence databases. AbSLang-search is highly compute efficient and allows to search datasets with 10 million sequences in less than 2 seconds.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wang, E. J. D., Spoendlin, F. C., Greenshields-Watson, A., Taylor, C. R., Deane, C. M.. 2026-08-11. Fast retrieval of structurally similar antibodies from large sequence databases with AbSLang. https://doi.org/10.64898/2026.08.05.742817
Cite the original work for its findings. Save a collection to share your selection of sources.