bioRxiv · 10.64898/2025.12.01.691162
Improving nanobody structure prediction with self-distillation
Abstract
Nanobodies are increasingly attractive therapeutic and biotechnological molecules, yet accurate structure prediction of their highly variable H-CDR3 loops remains a central challenge for machine learning models. Here, we investigate whether nanobody-specific structure prediction can be improved through curated synthetic data strategies. We systematically evaluate different data augmentation regimes, including self-distillation from unlabelled VHH sequences. To ensure structural plausibility of synthetic training samples, we develop NanoKink, the first sequence-based classifier of kinked versus extended H-CDR3 conformations, and apply stringent filtering criteria for non-canonical disulfide bond placement and confor-mational accuracy. On a curated benchmark enriched for challenging nanobody features, we show that, for a fixed training compute budget, a nanobody-specific model trained with filtered synthetic data significantly improves over baseline models and NanobodyBuilder2, achieving lower mean H-CDR3 RMSD and fewer structural violations, while remaining competitive with AlphaFold3 at approximately two orders of magnitude lower per-structure inference time. Our results highlight promising directions in synthetic data generation for nanobody structure modelling and provide a practical framework for optimisation of VHH structure prediction models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ali, M., Greenig, M., Jaskolowski, M., Crnogaj, M., Smorodina, E., Zhao, H., Greiff, V., Sormanni, P.. 2025-12-02. Improving nanobody structure prediction with self-distillation. https://doi.org/10.64898/2025.12.01.691162
Cite the original work for its findings. Save a collection to share your selection of sources.