Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.03.26.714386

Evolutionary history of ligand binding by the LRR domain of innate immunity receptors: the story of the TLR2 cavity

Abstract

Toll-like receptors (TLRs) are vital components of the innate immune system, recognizing both exogenous pathogens signals (PAMPs) and internal stress signals (DAMPs). TLR2 is unique among the human (Homo sapiens) TLR family members, as it contains a large cavity for binding hydrophobic ligands, such as lipoteichoic acid (LTA) and di/triacyl lipopeptides (Pam2/3CSK4). This study analyzed the structural phylogeny of cavity presence in the TLR2 lineage in vertebrates (vTLR) enabled by AI protein structure predictions and explored the potential convergent evolution of similar features in invertebrates (iTLRs). Analysis of AI models of TLR2s shows that this cavity is consistently present in TRL2 orthologs across jawed vertebrates (Gnathostomata). In jawless vertebrates (Cyclostomatha), these cavities were found in lamprey (Petromyzon marinus) TLR2 model, but only in some extant hagfish (Myxini), suggesting an ancestral origin in basal vertebrates followed by lineage-specific losses. TLR2 paralogs were found in several species, with a similar central cavity but potentially different ligand specificities. In silico ligand docking showed Pam2CSK4 binds to this cavity in all TLRs and paralogs consistently, demonstrating the conserved function of the ligand-binding pocket in gram-positive bacteria recognition across TLR2 branches. Changes in the TLR2 cavity size and shape in some vertebrate groups show the evolution of this DAMP recognition mechanism adapted to its respective pathogens. iTLRs form a separate phylogenetic branch with distinct structural features, but in literature some are considered to be TLR2 orthologs. Indeed, TLRs from some species of Helobdella and Ciona, contain a cavity with some similarity to that in the vTLR2 lineage. However, detailed structural comparisons of their location in the LRR domain and the structural details of the models suggest that their cavities have developed independently from that in TLR2s. Smaller cavities are present in other branches of the LRR family, but show different locations, shapes, and features, indicating that the binding of small ligands in the internal cavities within the LRR domains evolved multiple times in the LRR domain family history.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Namou, R., Ichii, K., Takkouche, A., Jaroszewski, L., Godzik, A.. 2026-03-30. Evolutionary history of ligand binding by the LRR domain of innate immunity receptors: the story of the TLR2 cavity. https://doi.org/10.64898/2026.03.26.714386

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A moving target: non-stationary selection governs unsupervised prediction of viral fitness

Anticipating how mutations change viral fitness is central to genomic surveillance and vaccine design, yet the supervised phenotype data behind the most accurate variant-effect predictors are unavailable for most emerging pathogens. We ask how far label-free scoring can go using only sequences, their evolutionary history, and structure. We assemble a modular, fully unsupervised pipeline that estimates a few interpretable terms (intrinsic replicative fitness, antigenic escape, and realized growth), and that lets each term be produced by more than one estimator, so the estimator itself becomes a testable modeling choice. Benchmarking the intrinsic term on 21 viral deep-mutational-scanning assays from ProteinGym, we find that a 650-million-parameter single-sequence protein language model predicts viral mutational fitness weakly and heterogeneously (mean Spearman 0.15), whereas a trivial site-independent alignment model more than doubles it (0.39, better on 17 of 21 assays), with the largest gains on the antigenic surface proteins where the language model fails. Yet the ordering reverses across 186 non-viral ProteinGym assays, where the language model instead exceeds the alignment model, localizing the weakness to viral families under-represented in the model's training data. Alignment-conditioned language models (MSA Transformer, Tranception) recover this accuracy but do not clearly exceed the simple alignment, so the decisive feature is the family alignment, not model scale or architecture. Our central result is evolutionary. Using dated samples of SARS-CoV-2 spike and influenza H3N2 hemagglutinin, we show that the epoch of the alignment is itself a leading, virus-specific determinant of accuracy. This traces to non-stationary selection: the site-specific amino-acid preferences drift over time, abruptly for spike at the emergence of Omicron and gradually for H3N2 hemagglutinin. A phylogenetic mutation-selection estimator does not match the far cheaper alignment model, falling significantly below it on matched data. Unsupervised viral fitness prediction is, then, as much an evolutionary problem as a modeling one.

bioinformatics↗

Accurate and scalable demultiplexing of single-cell RNA sequencing using BEACON

Barcode-based multiplexing strategies can significantly increase sample throughput and decrease costs, while mitigating batch effects of single-cell RNA sequencing experiments. However, these approaches can be limited by inaccurate or inefficient demultiplexing, resulting in cell loss and reduced statistical power. Here, we present BEACON, a novel sample demultiplexing method that efficiently learns the background count distribution from data and removes it from individual cells, thereby improving classification accuracy. BEACON outperforms other state-of-the-art methods on multiple human data sets. We apply it to cancer cell line time course experiments in vitro, enabling the identification of genes associated with aggressive tumors in vivo. Finally, we adapt BEACON to multimodal protein-transcriptome profiling, enhancing protein signal recovery to identify a CD161-positive effector memory CD4 T-cell population with a Th17-like phenotype, which we prospectively validate. BEACON can therefore be applied to other droplet-based single-cell sequencing methodologies.

bioinformatics↗

Gene-family-dependent thermodynamic effects of oncogenic mutations on nucleosome-DNA binding stability: a comparative molecular dynamics study across 22 cancer hotspots

Oncogenic mutations are known to occur at non-random rates and specific hotspots, yet the structural and thermodynamic factors underlying these hotspots remain poorly understood. This study presents molecular dynamics simulations of 22 cancer hotspots across 10 oncogene families, simulated as histone-DNA complexes. Across 18 of 22 mutations, thermodynamic effects were predominantly localized to within 6 [A] of the DNA-histone interface (mean capture 104%), with van der Waals interactions driving the effect at the thermodynamic extremes, indicating that oncogenic mutations alter nucleosome binding through precise local contact changes rather than global structural rearrangements. Further, analysis revealed a gene-family correlated pattern in thermodynamic stability. RAS family mutations showed a consistent trend toward nucleosome stabilization (mean {Delta}{Delta}G = -4.82 kcal mol-1), while kinase domain mutations trended toward destabilization (mean {Delta}{Delta}G = +40.12 kcal mol-1), a difference reaching statistical significance in this exploratory analysis (Mann-Whitney U, p = 0.008). This interface-specific mechanism, combined with the gene-family-correlated thermodynamic pattern, provides a biophysical framework for understanding nucleosome-level contributions to cancer hotspot mutation biology.

bioinformatics↗