bioRxiv · 10.64898/2026.05.06.723362
A fine-tuned genomic language model adds complementary nucleotide-context information to missense variant interpretation
Abstract
Missense variant interpretation remains a major challenge in clinical genomics. Existing missense variant impact predictors achieve strong performance, but they emphasize protein-level consequences and often share overlapping annotation priors. A missense annotation specifies the encoded amino-acid substitution, but the underlying nucleotide change may also act through nucleotide sequence context that protein-centric predictors overlook. Whether genomic language models capture distinctive nucleotide-level information beyond established missense variant impact predictors remains unclear. Through a comprehensive comparison of model backbones, embedding aggregation strategies, classifier heads, and adaptation regimes, we developed GLM-Missense, a genomic language model fine-tuned for missense variant impact prediction. Variant-position embeddings, multi-species pretraining, and low-rank adaptation were the design choices most critical to its performance. GLM-Missense contributed information complementary to established missense variant impact predictors. It showed low concordance with AlphaMissense, ESM1b, REVEL, CADD, SIFT and PolyPhen-2. We then asked whether this divergence was predictive rather than noise: after accounting for the other predictors, GLM-Missense retained the strongest unique association with variant pathogenicity. Finally, in MetaMissense, an XGBoost ensemble of all seven predictors, GLM-Missense ranked among the most informative features, indicating that its nucleotide-context signal carries information the ensemble uses to predict pathogenicity. A subset of the variants GLM-Missense resolved were annotated as missense but carried ClinVar splicing-related evidence for pathogenicity. This suggests GLM-Missense captures splicing-relevant signal. However, substituting SpliceAI for GLM-Missense in the ensemble failed to recapitulate its feature importance, suggesting that GLM-Missense captures additional sequence-derived information beyond splicing context alone. Together, these results demonstrate that GLM-Missense, a fine-tuned genomic language model, captures nucleotide-level information overlooked by established missense variant impact predictors.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Su, Y., Lin, Y.-J.. 2026-05-11. A fine-tuned genomic language model adds complementary nucleotide-context information to missense variant interpretation. https://doi.org/10.64898/2026.05.06.723362
Cite the original work for its findings. Save a collection to share your selection of sources.