bioRxiv · 10.64898/2026.02.16.706213
Guided tokenization and domain knowledge enhance genomic language models' performance
Abstract
Adapting language models to genomic and metagenomic sequences presents unique challenges, particularly in tokenization and task-specific generalization. Standard methods, such as fixed-length k-mers or byte pair encoding, often fail to preserve biologically meaningful patterns essential for downstream tasks. We introduce Guided Tokenization (GT), a strategy that prioritizes biologically and statistically important subsequences based on importance scores, model attention, and class distributions. Combined with domain adaptation, which incorporates prior domain specific biological knowledge, this approach improves both representation quality and classification accuracy in compact genomic language models (gLMs). GT enhances biological awareness in genomic language models, particularly for effective small and mid-sized models across key tasks, including DNA sequence read classification, promoter detection, antimicrobial resistance classification, and targeted amplicon taxonomic profiling. Our results highlight the promise of guided tokenization and domain-aware modeling for building efficient, biologically grounded language models for scalable genomic applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mahangade, V., Mollerus, M., Crandall, K. A., Rahnavard, A.. 2026-02-18. Guided tokenization and domain knowledge enhance genomic language models' performance. https://doi.org/10.64898/2026.02.16.706213
Cite the original work for its findings. Save a collection to share your selection of sources.