Search bioRxiv⌕ Search

Biology subjects

Shadskiy, A.

Publications and source records attributed to Shadskiy, A..

3 recordsLinked to original sources

GENATATORs: ab initio Gene Annotation With DNA Language Models

Inference of gene structure and location from genome sequences - known as de novo gene annotation - is a fundamental task in biological research. However, sequence grammar encoding gene structure is complex and poorly understood, often requiring costly transcriptomic data for accurate gene annotation. In this work, we benchmark current solutions and develop new methods of gene annotation. We show that pre-trained DNA language model (DNA LM) embeddings do not capture the features necessary for precise gene segmentation, and that task-specific fine-tuning remains essential. We comprehensively evaluate the impact of model architecture, training strategy, receptive field size, dataset composition, and data augmentations on gene segmentation performance. We revisit standard evaluation protocols, showing that commonly used per-token and per-sequence metrics fail to capture the challenges of real-world gene annotation. We introduce and theoretically justify new biologically grounded metrics, along with benchmarking datasets that better capture annotation quality. We show that fine-tuned DNA LMs outperform existing annotation tools, generalizing across species separated by hundreds of millions of years from those seen during training, and providing segmentation of previously intractable non-coding transcripts and untranslated regions of protein-coding genes. Our results thus provide a foundation for new biological applications centered on accurate gene annotation.

bioinformatics↗

Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for DNA Foundation Models

AO_SCPLOWBSTRACTC_SCPLOWRecent advances in DNA language models have mainly come from building larger and more complex architectures, making it harder to understand the effect of changes to standard components such as the transformer layers widely used in NLP. In this work, we study whether and how a modernized BERT-style back-bone (ModernBERT) can be adapted to genomic sequence modeling to improve computational efficiency, training stability, and downstream performance. Under controlled experimental settings, we benchmark efficiency across a range of sequence lengths and evaluate downstream performance on the Nucleotide Transformer benchmark. The resulting model, ModernGENA, achieves a strong efficiency-quality trade-off and ranks among the top-performing models in our evaluation suite. To support reproducibility and provide a solid default reference point for future architectural work in genomics, we release the full implementation and configuration of ModernGENA as an open, reusable baseline, and make ModernGENA base and ModernGENA large publicly available through the DNA language models collection on Hugging Face.

bioinformatics↗

Expanding the list of sequence-agnostic enzymes for chromatin conformation capture assays with S1 nuclease

This study presents a novel approach for mapping global chromatin interactions using S1 nuclease, a sequence-agnostic enzyme. We develop and outline a protocol that leverages S1 nucleases ability to effectively introduce breaks into both open and closed chromatin regions, allowing for comprehensive profiling of chromatin properties. Our S1 Hi-C method enables the preparation of high-quality Hi-C libraries, marking a significant advancement over previously established DNase I Hi-C protocols. Moreover, S1 nucleases capability to fragment chromatin to mono-nucleosomes suggests the potential for mapping the three-dimensional organization of the genome at high resolution. This methodology holds promise for an improved understanding of chromatin state-dependent activities and may facilitate the development of new genomic methods.

molecular biology↗