Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.01.17.700093

Fundamental limitations of genomic language models for realistic sequence generation

Abstract

Large language models (LLMs) have shown remarkable success in natural language processing, prompting interest in their application to genomic sequence analysis. Genomic Language Models (gLMs) based on similar architectures offer a promising avenue for synthetic genome generation and characterization. However, their effectiveness for biological sequence modeling remains poorly characterized. We present a comprehensive evaluation of genomic language models that explicitly aim to generate entire synthetic genomes. We tested Evo 2 on diverse prokaryotic, eukaryotic and viral genomes, and megaDNA on bacteriophage genomes, and assessed performance across key biological features and organizational patterns. Our results reveal systematic failures in gLM-based genomic reconstruction. While the synthetic sequences captured local sequence statistics, they consistently failed to preserve long-range genomic organization, repeat and k-mer composition, transcription factor binding site architecture, and evolutionary constraints. Generated sequences exhibited violations of natural genomic patterns and models showed particular difficulty with repetitive elements. To assess the quality of genome generation, we trained a convolutional neural network that reliably distinguished synthetic from natural sequences, achieving AUROC values up to 0.97 in eukaryotes and 0.82 in prokaryotes, with classification accuracy increasing monotonically with genomic distance from the seed. These findings suggest fundamental limitations in current gLM architectures for capturing the long-range, hierarchical nature of genomic sequences. Our work highlights the need for specialized architectures that explicitly model evolutionary constraints rather than relying solely on statistical patterns, with important implications for computational biology applications requiring realistic sequence generation and for biosafety assessments that depend on the distinguishability of synthetic and natural genomic sequences.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tzanakakis, A., Mouratidis, I., Georgakopoulos-Soares, I.. 2026-01-18. Fundamental limitations of genomic language models for realistic sequence generation. https://doi.org/10.64898/2026.01.17.700093

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Data-driven predictive design of engineered living hydrogels

Engineered living materials (ELMs) offer a promising route to biologically manufactured materials for healthcare, construction and manufacturing. However, their rational design is limited by the lack of quantitative relationships linking design parameters to material properties. Here, we show that a tabular foundation model (TabPFN), informed by a small library of living hydrogels, accurately predicts macroscopic material properties from genetic and process parameters. Models were evaluated for predicting storage modulus (G'), fibre content, thickness, and permeability of Escherichia coli-produced living hydrogels containing CsgA-based fibres fused to genetically encoded PEG-like biopolymers. On an independent validation set, TabPFN achieved the strongest prediction for G' (R2 = 85.1%), reducing RMSE by 48.0% compared to linear regression. Property-guided design further enabled identification of parameters for achieving living hydrogels with desired properties. These results establish a broadly applicable framework for predictive design of ELMs, reducing experimental screening and accelerating the discovery of materials with targeted properties.

synthetic biology↗

Characterization of a fungal mixed-linkage glucan synthase

Mixed-linkage (1,3;1,4)-{beta}-glucan (MLG) is a cell wall polysaccharide found in fungi, bacteria, and plants, yet the enzymatic machinery and structural determinants governing fungal MLG biosynthesis remain poorly understood. In Aspergillus fumigatus, TFT1 (AfTFT1) has previously been implicated in MLG biosynthesis, but its direct biochemical activity remained unresolved. The structural characterization of A. fumigatus wall MLG revealed a distinct polymer profile. Heterologous expression of AfTFT1 in Komagataella phaffii enabled MLG production, generating an oligosaccharide profile that closely matched that of the native fungal polymer and demonstrating that AfTFT1 functions as an MLG synthase. Phylogenetic analysis placed AfTFT1 within a distinct glycosyltransferase family 2 (GT2) fungal lineage containing related candidate MLG synthases across diverse filamentous Ascomycota and separate from the major plant and bacterial synthase lineages. Structure-guided mutagenesis showed that individual substitutions within the transmembrane pore and switch motif altered the lichenase-derived oligosaccharide profile while retaining detectable MLG production, whereas replacement of the entire switch motif with the corresponding HvCSLF6 sequence resulted in no detectable MLG production. Together, these findings establish AfTFT1 as a fungal MLG synthase and identify the switch motif and adjacent transmembrane region as important determinants of MLG synthase function and product fine structure, providing insight into the structural and evolutionary diversification of MLG biosynthesis.

synthetic biology↗

Firewalled synthetic commensal blocks horizontal gene transfer in the gut

Synthetic biology enables the rational reprogramming of microorganisms into living therapeutics and agents for bioremediation. However, such genetically modified organisms (GMOs) disseminate their synthetic genetic information into natural microbial communities through horizontal gene transfer (HGT), posing biosafety risks that limit clinical and environmental deployment. Reassigning sense codons to an alternative amino acid identity establishes a genetic firewall that simultaneously prevents incoming and outgoing gene flow, but reported implementations compromise fitness, precluding clinical and industrial use. Here, we overcome this limitation using genome design and laboratory evolution to create a high-fitness genetically firewalled Escherichia coli commensal. By directly altering the amino acid identity of TCA and TCG serine codons in the genetic code without an unassigned intermediate, we establish a robust genetic firewall that remains stable for thousands of generations. This firewalled commensal stably colonizes the mouse gastrointestinal tract for more than 100 days and blocks viral infections and HGT. As the long-term within-gut evolution of this firewalled organism identified adaptive mutations in genes responsible for carbon source utilization, we rationally redesigned the strain's genome to increase fitness. Together, this work establishes a genetically firewalled commensal for safer living therapeutics development and provides a strategy for designing high-fitness, virus- and gene-transfer-resistant organisms for clinical and environmental use.

synthetic biology↗