bioRxiv · 10.64898/2026.09.02.749017
EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
Abstract
Hybrid DNA foundation models combine convolutional, recurrent, and attention layers, making speculative decoding more difficult than truncating a KV cache. We present EvSpark, a speculative decoding system for Evo2 that verifies draft blocks in parallel and restores all three classes of inference state by selecting retained intermediate states, without replay. A compact, hidden-state-conditioned drafter proposes each block in one parallel forward pass. On Evo2 7B, a 48-prompt benchmark with three training seeds yields 2.96x on 43 real-sequence prompts and 3.27x including the five synthetic controls. Acceleration persists at 262k-token context (1.84x - 2.43x on two bacterial genomes) and over 32k generated tokens. Retraining the same drafter architecture for Evo2 20B and 40B yields 2.18x - 2.46x on real sequences and 2.51x - 2.78x on the full suite. Autoregressive drafter comparisons and batch measurements show why low draft latency, rather than acceptance alone, determines the gain. The method preserves the target distribution in exact arithmetic. In bf16, greedy tests find no non-tie divergences across 48 prompts and six checkpoints; sampling tests expose residual numerical sensitivity, especially in repetitive sequences. A cost-efficient 7B drafter requires 1.06 incremental GPU-hours of training, excluding teacher-data collection, and achieves 2.82x on real sequences. In regulatory-DNA design, EvSpark achieves a median complete-workflow speedup of 1.57x over a calibrated batched native baseline.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ding, H., Wu, N., Qiu, T.. 2026-09-08. EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models. https://doi.org/10.64898/2026.09.02.749017
Cite the original work for its findings. Save a collection to share your selection of sources.