bioRxiv · 10.64898/2026.06.30.735115
CodonBERT and ESM-2 Embedding Spaces Share an Evolutionarily Conserved Paired Geometry Encoding Synonymous Codon Information
Abstract
1.1Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of in-dependently trained deep learning models--one operating on coding DNA sequences (CDS) and the other on protein sequences--remains an open question. Here we systematically examine the paired geometry between two embedding spaces: CodonBERT, a nucleotide language model trained exclusively on CDS with codon-aware tokenization, and ESM-2, a protein language model. After removing linear effects of amino acid composition, we find that CodonBERT and ESM-2 embeddings exhibit a robust, linearly retriev-able paired correspondence across human and mouse transcriptomes (R@1 = 0.266, 95% CI [0.256, 0.276]), with a synonymous-codon-specific signal fraction of {Delta}R@1 {approx} 0.08. This paired geometry transfers faithfully between species and is sharply amplified under ortholog restriction (R@1 = 0.461-0.538). Critically, the per-gene alignment cosine--measured in a gallery-free framework that avoids retrieval-set-size artifacts--decays monotonically with evolutionary distance across vertebrates (median cosine: rat 0.52 [->] mouse 0.47 [->] zebrafish 0.44 [->] yeast 0.19; all adjacent comparisons p < 10 {superscript 2}, Mann-Whitney U), and a shuffled-pair permutation null confirms that the observed alignment is absent when biological pairing is destroyed (null mean {approx} 0.008; p < 10 ). Ablation experiments confirm that this signal is specific to codon-aware nucleotide architectures: CodonBERT substantially outperforms DNABERT-2 and Nucleotide Transformer v2 in per-gene alignment cosine (median 0.47 vs. 0.28 vs. 0.26; Mann-Whitney p < 10 ). These results demonstrate that synonymous-codon-level regulatory information is embedded in an evolutionarily constrained geomet-ric relationship between coding sequence and protein representations, which can be linearly decoded without any joint training.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lu, B.. 2026-07-03. CodonBERT and ESM-2 Embedding Spaces Share an Evolutionarily Conserved Paired Geometry Encoding Synonymous Codon Information. https://doi.org/10.64898/2026.06.30.735115
Cite the original work for its findings. Save a collection to share your selection of sources.