Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.07.09.737627

Semantic fragment representations for coordinate-free analysis of genomics data

Abstract

Many genomic assays begin with individual DNA fragments, but standard analysis quickly collapses those molecules into counts over genomic intervals. Rich information carried by each fragment, including its sequence, fragment body, cleavage boundaries, and local flanking context, is lost in this process. This loss is especially apparent in mixed-source and heterogeneous samples, where individual fragments originate from disparate cell types and can retain information about their cell of origin. To address this, we present LEAF-1, a fragment-level foundation model pre-trained on approximately 58 billion fragments spanning bulk ATAC-seq, single-cell ATAC-seq, and cell-free DNA profiles, representing each DNA molecule as a point in a learned semantic space defined by sequence context, assay modality, and explicit cleavage-boundary tokens. In sparse scATAC-seq datasets, mean-pooled LEAF-1 embeddings readily classify human cell types from as few as [~]1,000 fragments per cell, with high-scoring fragments linked to cell-type-associated transcription-factor programs. Similarly, in cell-free DNA profiling, LEAF-1 outperformed state-of-the-art coordinate-binning strategies and general-purpose DNA language model baselines across cancer detection tasks. Applying attention-based multiple-instance learning to LEAF-1 embeddings further improved cancer detection, reaching an area under the receiver operating characteristic (ROC) curve (AUC) of 0.95. This pan-cancer model generalizes beyond cancer types it is trained on, as we show by profiling plasma samples from clear cell renal cell carcinoma patients and healthy volunteers and applying the frozen classifier without retraining, achieving an AUC of 0.83. These results show that semantic learning over individual DNA fragments preserves biochemical, cell-associated, and disease-associated signals that are otherwise lost during coordinate-based aggregation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Heydari, H., Zhao, J., Arseneault, M., Younesian, L., Tanguay, S., Riazalhosseini, Y., Goodarzi, H., Najafabadi, H. S.. 2026-07-10. Semantic fragment representations for coordinate-free analysis of genomics data. https://doi.org/10.64898/2026.07.09.737627

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Interrogation of noncoding schizophrenia risk variants using CRISPR-based functional genomics

Schizophrenia (SCZ) is a highly heritable complex disorder influenced by coding and noncoding genetic variation. Its genetic causes, particularly those involving noncoding variation, are largely unknown. High-throughput CRISPR screens enable dissection of disease-associated loci and identification of noncoding regulatory elements and variants that modulate gene expression. We screened SCZ GWAS loci linked to genes that are also associated in whole-exome sequencing studies to identify regulatory elements and variants impacting expression of disease-relevant genes. We used CRISPRi paired with HCR-FlowFISH to epigenetically silence 333 putative regulatory elements and measure the downstream effects on gene expression of causal SCZ genes, FAM120A, SV2A, and STAG1, in iPSCs and iPSC-derived neurons (iNeurons). We identified 78 regulatory elements that significantly alter expression of a SCZ gene, including noncoding enhancers/silencers as well as promoters of genes and lncRNAs. Pooled prime editing screens interrogated noncoding variant influence on gene expression for SCZ-associated variants and uncharacterized common variants from diverse population studies. We find that a common variant in the promoter of SV2A, rs112851681:A>G (MAF = 3.56%, 1000 Genomes) enhances transcriptional activity in iPSCs and iNeurons. These findings show distinct noncoding mechanisms that map within GWAS signals, and provide a path forward for interrogating noncoding regulatory elements and variants in disease loci.

genomics↗

The automated eukaryotic pangenome pipeline EukPan reveals accessory genome differentiation beyond core-gene phylogeny in Aspergillus oryzae

Pangenome analysis reveals recurrent gene-content variation beyond a single reference genome, but its application to eukaryotes is constrained by inconsistent gene annotation. ANNEVO predicts gene models from genome FASTA assemblies without RNA-seq data. We developed EukPan, an automated post-annotation pipeline that standardizes GFF/GTF files, selects representative isoforms, constructs proteomes, infers orthogroups, builds a concatenated single-copy core-protein alignment, and summarizes shared accessory orthogroups while excluding orthogroups detected in only one genome. Applied with ANNEVO to 123 Aspergillus oryzae genomes, EukPan identified 11,245 core and 4,407 shared accessory orthogroups. The core-protein phylogeny broadly recovered the reported A-H classification, whereas accessory-genome analyses clearly separated the 33 group-A strains from the other 90 strains. Directional analysis identified 62 group-A-associated and 158 group-A-depleted orthogroups, with major facilitator superfamily (MFS) transporter and fungal Zn2Cys6 transcription-factor domains prominent in the depleted set. Among 93 orthogroups present in all non-A strains and absent from all group-A strains, 59 mapped to 10 segments of RIB40, the standard A. oryzae reference genome and a non-A (group-F) strain. EukPan therefore enables reproducible, coordinated core- and accessory-pangenome analysis from eukaryotic genome assemblies.

genomics↗

Multi-Omics analysis provides crucial insights into ecological adaptation to dryland of a dominant grass (Psammochloa villosa, Poaceae) in Northwest China

Desertification exerts dramatic selection pressures on the evolution of plants. Despite the key role of ecological adaptation by natural selection to arid grasslands and subsequent intraspecific divergence, specific mechanisms driving this process remain poorly understood. Psammochloa villosa, a perennial forage grass endemic to the arid grasslands in Northwest China, where it thrives in shifting and semi-fixed sand land due to its exceptional drought tolerance, provides an ideal system to study adaptive evolution to aridity. In our study, we assembled a high-quality, chromosome-scale genome and conducted genomic resequencing of 42 populations across its major distribution. The genome assembly, which is approximately 1.55 Gb in size, has a super-scaffold N50 of 66.79 Mb, with 75.84% of the sequences identified as transposable elements. Coalescent phylogeny and genomic collinearity analyses strongly supported that P. villosa and Neotrinia splendens, as the closest taxa, shared a recent whole-genome duplication (WGD) event occurring approximately 18-20 Mya and followed by their divergence around ~11.2 Mya. Based on ancestral grass karyotype (AGK) reconstruction from synteny analysis, our results suggest that, relative to the AGK after the {rho}-WGD event, P. villosa and N. splendens underwent similar chromosomal restructuring and lineage-specific retention of numerous copies, providing a genomic basis for potential ecological adaptation and intraspecific diversification. The expanded XTH family, which encodes enzymes mediating xyloglucan endotransglucosylation and hydrolysis and thereby regulating xyloglucan remodeling, showed strong transcriptional responses under PEG-6000 treatment, suggesting that retained copies may be associated with xerophytic adaptation in P. villosa. Together, these findings suggest that WGD-derived gene retention created a delayed reservoir of genetic diversity that was later shaped by desertification and Qinghai-Xizang Plateau environmental changes, contributing to climate-associated genomic islands and intraspecific differentiation.

genomics↗