bioRxiv · 10.64898/2026.09.21.753204
Limitations of general protein language models and public TCRpMHC data in training specificity prediction models
Abstract
T cell receptor (TCR) specificity prediction is critical for understanding adaptive immunity and for enabling the development of therapeutic interventions. Current machine learning models for predicting TCR-peptide-MHC interactions perform poorly, despite large efforts by many groups and advances in protein language models. This work identifies and addresses the fundamental issues that might be hindering progress in this field, focusing on current approaches for constructing training and validation sets, particularly on the strategies used by commonly used methods, which generate negative data either by shuffling TCRs between epitopes or by pairing epitopes with TCRs from background datasets of healthy donors. Here, we visualize and quantify the overlap between positive and generated negative TCR sets produced by both strategies using a range of sequence embedding approaches, including simple encodings (one-hot and BLOSUM62), general protein language models (ESM-2 and ProtGPT2), and a TCR-specific embedding model (TCR-BERT). Using paired {beta}TCR data from VDJdb for two well-studied epitopes (GILGFVFTL and KLGGALQAK), we train Random Forest classifiers on each embedding representation and evaluate their ability to distinguish positive from negative data examples. We find that general sequence embedding methods struggle to differentiate between positive and negative TCRs. We further show that the choice of negative data generation method significantly impacts classifier performance. Our findings highlight fundamental limitations in current TCR specificity training data and general protein language models, and underscore the need for TCR-specific embeddings, experimentally validated negative datasets, transparent reporting of negative data construction, and epitope-specific modeling approaches.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rai, B., Hall-Swan, S.. 2026-09-25. Limitations of general protein language models and public TCRpMHC data in training specificity prediction models. https://doi.org/10.64898/2026.09.21.753204
Cite the original work for its findings. Save a collection to share your selection of sources.