bioRxiv · 10.64898/2026.09.08.750172
Causal variant underestimation is a major overlooked driver of sequence-to-function model underperformance
Abstract
Deep learning sequence-to-function (S2F) models represent tremendous promise for functionally fine-mapping causal variants associated with traits and disease. Yet it remains unclear precisely how effective they are at this task. Generally, S2F models perform well at classifying putatively causal expression quantitative trait loci (eQTL) SNVs, yet dramatically underperform linear baselines at ranking different individuals' gene expression values from their whole genome sequence, which directly calls into question their ability to fine-map a locus by properly weighting the effects of variants onto gene expression. Here, using the state-of-the-art S2F model AlphaGenome, we systematically compared predicted effect sizes for fine-mapped eQTLs to nearby putatively non-causal SNVs, finding that AlphaGenome pervasively underestimates the effects of most causal variants and fails to successfully fine-map eQTLs in most loci for this reason. We show that misdirected variant effect predictions and negative cross-individual correlations, widely cited as major challenges facing S2F expression modeling, are not egregious errors but a chance consequence of weak, noisy attributions when causal variants are underestimated. Our results suggest that a failure to detect local variant effects onto enhancer activity is a cause of underestimation. Additionally, we find that underestimated variants are enriched far from their target gene, suggesting inadequate enhancer-gene linking causes underestimation even when local effects are well-detected. Finally, our results clarify when S2F models are reliable for fine-mapping objectives: while they pervasively underestimate most causal variants, demonstrating limited fine-mapping utility for most loci, the top 0.1 most prominent variant effect predictions are strongly enriched for causal variants with directionally correct predictions that are consistent across model replicates, suggesting extremely strong predictions are broadly trustworthy. Our results demonstrate that causal variant underestimation is a core issue facing S2F expression predictors, with future improvements dependent on better local activity detection and enhancer-gene linking.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Drusinsky, S., Pollard, K. S.. 2026-09-11. Causal variant underestimation is a major overlooked driver of sequence-to-function model underperformance. https://doi.org/10.64898/2026.09.08.750172
Cite the original work for its findings. Save a collection to share your selection of sources.