Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.07.10.737788

What Do Generative Models Learn About Adaptive Immune Receptor Repertoires? A Benchmark Study

Abstract

Generative models are increasingly used to model adaptive immune receptor repertoire (AIRR) sequence distributions, promising to decode the sequence diversity shaping immune responses and accelerate the design of therapeutic antibodies and T-cell receptors. Yet it remains unclear whether these models produce biologically meaningful outputs or merely capture surface-level sequence statistics while missing features driven by receptor generation and selection. Rigorous evaluation is needed, but the field lacks established standards, as existing machine learning metrics do not all translate directly to the AIRR domain, given the complex structure of the data and the lack of biological ground truth. Consequently, researchers face difficulties in evaluating the models and selecting appropriate ones, which can critically affect downstream clinical applications. Here, we apply a suite of evaluation metrics tailored to AIRR sequence data and present a systematic comparison of popular generative model families proposed for the AIRR field, including variational autoencoders, long short-term memory networks, antibody language models, selection models, and simple statistical baselines. We focus specifically on the task of learning individual-specific immune receptor repertoires, a clinically relevant challenge with direct implications for personalized immunotherapy, disease monitoring, and vaccine response studies. By analyzing the sequences generated by each model, we identify memorization risks, innovation capabilities, and sensitivity to hyperparameter tuning. Taken together, these results advance the understanding of how current generative models reproduce the biology of individual immune repertoires and lay the groundwork for more principled model development and evaluation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Würtzen, C., Mamica, M., Kanduri, C., Pavlovic, M., Greiff, V., Peters, B., Sandve, G. K.. 2026-07-16. What Do Generative Models Learn About Adaptive Immune Receptor Repertoires? A Benchmark Study. https://doi.org/10.64898/2026.07.10.737788

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Melanoma suppresses galectin-9-glycan axis in dendritic cells and galectin-9 restoration limits T regulatory cell expansion

Dendritic cells (DCs) are key orchestrators of anti-tumor adaptive immune responses. Their function is tightly regulated by galectins, a family of carbohydrate-biding proteins that decode extracellular glycans into intracellular signaling. Here, we show that exposure to melanoma-conditioned media (CM) induces loss of galectin-9 (gal-9) at the DC surface. Notably, gal-9 loss, independent of transcriptional regulation, was associated with the acquisition of an immunosuppressive phenotype in CD14cDC2 cells following tumor exposure, with higher-molecular weight fractions of the melanoma secretome as mediators. Notably, gal-9 depletion was mirrored with a reduction in gal-9 ligands on the cell surface, particularly GalNAc-containing glycoepitopes, pointing towards a melanoma-exploited gal-9-glycan axis as a DC-evasive strategy. Interestingly, restoring gal-9 surface levels in CM-exposed CD14cDC2 cells prevented the expansion of regulatory T cells (Tregs), postulating gal-9 as a novel immunomodulatory molecule in DC-mediated Treg induction during melanoma progression. Altogether, our data suggest that melanoma-derived factors remodel the DC glycan-gal-9 axis to enhance T cell differentiation towards regulatory phenotypes and dampen anti-tumor immunity. This identifies the gal-9/glycan axis in CD14cDC2 cells as a vulnerable node in melanoma immune evasion and as a potential therapeutic target.

immunology↗

Repeated shrimp allergen exposure drives 5-lipoxygenase-dependent avoidance and selective gut-brain activation

Peripheral immune processes can shape animal behavior, yet how noninfectious inflammatory reactions affect neural activity and behavioral outputs remains poorly understood. We developed an optimized murine model of shrimp allergy using whole shrimp extract to examine how a complex dietary allergen elicits integrated immune, neural, and behavioral responses. Sensitized mice received repeated oral shrimp challenges and were assessed for allergic pathology, food preference, affective-like behaviors, and neuronal activation in the brain. Repeated exposure increased total IgE and shrimp-specific IgG1, induced mast cell activation, accelerated gastrointestinal transit, caused mild hypothermia consistent with oral anaphylaxis, and increased intestinal length. Shrimp-sensitized mice did not avoid shrimp solution after sensitization alone. Instead, avoidance emerged only after repeated oral challenges and strengthened over time. This delayed aversion occurred without detectable changes in locomotor activity or measures of anxiety-like or depressive-like behavior at the time points tested. Repeated shrimp exposure increased cFOS expression in the area postrema, nucleus of the tractus solitarius, central amygdala, and paraventricular nucleus of the thalamus, implicating brainstem and limbic-thalamic pathways involved in visceral sensing and aversion. Pharmacological inhibition of 5-lipoxygenase partially reversed avoidance and reduced circulating mast cell protease-1 in allergic mice. These findings establish a robust whole-shrimp allergy model and show that a complex food allergen engages gut-brain pathways to promote 5-lipoxygenase-dependent avoidance. The delayed, selective nature of this response supports immune-mediated food aversion as a shared output of food allergy, while suggesting that its kinetics and neural recruitment vary with allergen identity and inflammatory context.

immunology↗

Targeting IL-2 to inflamed tissues via oxidation-specific epitopes enables third-generation bispecific IL-2 therapeutics

Interleukin-2 (IL-2) is essential for the survival and activation of regulatory T cells (Tregs). Low-dose native IL-2 (IL-2LD) therapy restores immune regulation in vivo and has shown reproducible clinical benefit across multiple autoimmune, inflammatory, and neuroimmune diseases. Attempts to improve IL-2 through engineered variants (muteins) have mainly focused on enhancing Treg selectivity by reducing IL-2 receptor {beta}-chain binding, but this strategy profoundly diminishes biological potency, likely contributing to the limited clinical efficacy of IL-2 muteins. Here, we develop a ''third-generation IL-2'' that combines site-specific targeting and bifunctionality. We generated a bivalent fusion protein linking IL-2 to a single-chain antibody recognizing oxidation-specific epitopes (OSEs), which are abundantly expressed at inflamed sites. Targeting OSEs provides not only site-specific localization, but also true bifunctionality as both anti-OSE antibodies and IL-2LD independently show therapeutic benefit in limiting inflammation. We first show that IL-2IT has bifunctional biological activities in vitro. In vivo, IL-2IT had increased specificity for Treg over Teff activation, which we attribute to a conformation-dependent modulation of IL-2 receptor engagement. Importantly, IL-2IT provided precise delivery to inflamed tissues in models of psoriasis and colitis. Altogether, this resulted in superior therapeutic benefit in multiple clinical settings, including in atherosclerosis models. Thus, our strategy illustrates a generalizable approach to cytokine engineering that preserves native signaling while achieving spatial control. Specifically, our findings validate OSE targeting as an efficient strategy to guide therapeutics to sites of inflammation and establish OSE-IL-2 as a promising bispecific Treg engager for treating inflammation.

immunology↗