bioRxiv · 10.64898/2026.02.05.703637
Short-Context Regulatory DNA Language Models with Motif-Discovery Regularization
Abstract
DNA language models (DNALMs) aim to learn representations of genomic sequence for variant interpretation, regulatory prediction, and sequence design. Most DNALMs are trained on whole genomes and long contexts, but regulatory DNA poses a distinct challenge: functional elements are sparse, context dependent, and encoded by short transcription factor motif syntax embedded in extensive background sequence. We introduce ARSENAL, a short-context masked DNA language model pretrained on ENCODE candidate cis-regulatory elements. ARSENAL recovers diverse transcription factor motifs de novo and improves zero-shot regulatory variant effect prediction relative to other DNALM foundation models. ARSENAL embeddings also improve supervised regulatory sequence models at predicting chromatin accessibility and regulatory variant scoring. Finally, ARSENAL serves as an efficient generative prior, enabling multi-objective regulatory sequence design with supervised oracles. ARSENAL shows that targeted self-supervised pretraining on regulatory regions can learn biologically meaningful and transferable regulatory representations without genome-scale training, long contexts or task-specific labels. Code is available at https://github.com/kundajelab/regulatory_lm. Models and data are shared at https://sageb.io/ydjhqM
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Patel, A., Kundaje, A.. 2026-02-06. Short-Context Regulatory DNA Language Models with Motif-Discovery Regularization. https://doi.org/10.64898/2026.02.05.703637
Cite the original work for its findings. Save a collection to share your selection of sources.