Enhancing CAR-T cell activity prediction via fine-tuning protein language models with generated CAR sequences
Chimeric antigen receptor (CAR)-T cell therapy has shown remarkable success in treating hematological malignancies; however, several challenges remain, including limited efficacy against solid tumors, T cell exhaustion, and lack of T cell persistence, which have restricted its clinical efficacy across various indications. Sequence optimization of CAR constructs offers a promising strategy to enhance therapeutic efficacy of CAR-T cells. Recent advances in machine learning, especially protein language models (PLMs), enable prediction of mutational effects based on sequence representations. Nevertheless, applying PLMs to CARs is challenging due to the artificial nature of CARs and the absence of comprehensive CAR sequence databases. In this study, we developed a computational framework to predict CAR-T cell activity by fine-tuning ESM-2 with the CAR sequences generated using sequence augmentation. These CAR sequences were constructed by recombining homologous domains of CARs, enabling task-specific adaptation of the model. To evaluate prediction performance, we experimentally assessed the cytotoxicity of CAR-T cells expressing mutated CAR variants and compared these results with model predictions. Our results demonstrated that fine-tuned ESM-2 significantly improves prediction performance of CAR-T cell activity. Furthermore, we showed that training parameters--such as sequence diversity, number of training steps, and model size--substantially influence prediction performance. This work highlights the potential of combining sequence augmentation with fine-tuning PLMs to advance data-driven CAR-T cell design.