Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.03.05.641690

Integrative genotyping and analysis of canine structural variation using long-read and short-read data

Abstract

Structural variation makes an important contribution to canine evolution and phenotypic differences. Although recent advances in long-read sequencing have enabled the generation of multiple canine genome assemblies, most prior analyses of structural variation have relied on short read sequencing. To offer a more complete assessment of structural variation in canines, we performed an integrative analysis of structural variants present in 12 canine samples with available long-read and short-read sequencing data along with genome assemblies. Use of long-reads permits the discovery of heterozygous variation that is absent in existing haploid assembly representations while offering a marked increase in the ability to identify insertion variants relative to short-read approaches. Examination of the size spectrum of structural variants shows that dimorphic LINE-1 and SINE variants account for over 45% of all deletions and identified 1,410 LINE-1s with intact open reading frames that show presence-absence dimorphism. Using a graph-based approach, we genotype newly discovered structural variants in an existing collection of 1,879 resequenced dogs and wolves, generating a variant catalog containing a 56.5% increase in the number of deletions and 705% increase in the number of insertions previously found in the analyzed samples. Examination of allele frequencies across admixture components present across breed clades identified 283 structural variants evolving with a signature of selection. Significance statementExisting studies of structural variation have focused on genomes sequenced with short-read technology. In this study, we systematically combine short-read and long-read data sets to identify and genotype structural variation across canines, resulting in an expanded catalog of structural variants. Analysis of allele frequencies identified structural variants in genic regions that may have been evolving under selection across breed clades.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Schall, P. Z., Kidd, J. M.. 2025-03-11. Integrative genotyping and analysis of canine structural variation using long-read and short-read data. https://doi.org/10.1101/2025.03.05.641690

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

ssJSD: A fusion of sparsity and spatial information for HiC single-cell clustering

Single-cell high-throughput chromatin conformation capture (scHiC) enables profiling three dimensional genome architecture at cellular resolution, providing insights into cell-to-cell variability and cellular functions. Recent frameworks utilize spatial interaction patterns to derive dissimilarity measures for downstream tasks such as cell clustering. However, the inherent sparsity and ultra-high dimensionality of scHiC contact matrices pose significant challenges. A central hurdle is that existing measures typically treat all zeros without distinction, failing to differentiate biologically meaningful structural zeros (SZs) from technical dropouts. Here, we introduce ssJSD (spatial and sparsity informed Jensen-Shannon Divergence), a computational framework designed to explicitly account for scHiC-specific sparsity patterns. By integrating band-wise contact frequency profiles with SZ-induced sparsity matrices, ssJSD leverages both spatial interaction patterns and biological absence of contacts. We adopted two complementary integration strategies: an early fusion approach that concatenates information into a single representation, and a late fusion approach that integrates JSD-based dissimilarities through diverse averaging methods. Through simulations and applications to human cell lines and prefrontal cortex data, we demonstrate that ssJSD improves clustering accuracy and effectively distinguishes cell types. Our study indicates that integrating SZ patterns is important for accurately quantifying cell-to-cell variability in 3D genomics.

genomics↗

Interrogation of noncoding schizophrenia risk variants using CRISPR-based functional genomics

Schizophrenia (SCZ) is a highly heritable complex disorder influenced by coding and noncoding genetic variation. Its genetic causes, particularly those involving noncoding variation, are largely unknown. High-throughput CRISPR screens enable dissection of disease-associated loci and identification of noncoding regulatory elements and variants that modulate gene expression. We screened SCZ GWAS loci linked to genes that are also associated in whole-exome sequencing studies to identify regulatory elements and variants impacting expression of disease-relevant genes. We used CRISPRi paired with HCR-FlowFISH to epigenetically silence 333 putative regulatory elements and measure the downstream effects on gene expression of causal SCZ genes, FAM120A, SV2A, and STAG1, in iPSCs and iPSC-derived neurons (iNeurons). We identified 78 regulatory elements that significantly alter expression of a SCZ gene, including noncoding enhancers/silencers as well as promoters of genes and lncRNAs. Pooled prime editing screens interrogated noncoding variant influence on gene expression for SCZ-associated variants and uncharacterized common variants from diverse population studies. We find that a common variant in the promoter of SV2A, rs112851681:A>G (MAF = 3.56%, 1000 Genomes) enhances transcriptional activity in iPSCs and iNeurons. These findings show distinct noncoding mechanisms that map within GWAS signals, and provide a path forward for interrogating noncoding regulatory elements and variants in disease loci.

genomics↗

The automated eukaryotic pangenome pipeline EukPan reveals accessory genome differentiation beyond core-gene phylogeny in Aspergillus oryzae

Pangenome analysis reveals recurrent gene-content variation beyond a single reference genome, but its application to eukaryotes is constrained by inconsistent gene annotation. ANNEVO predicts gene models from genome FASTA assemblies without RNA-seq data. We developed EukPan, an automated post-annotation pipeline that standardizes GFF/GTF files, selects representative isoforms, constructs proteomes, infers orthogroups, builds a concatenated single-copy core-protein alignment, and summarizes shared accessory orthogroups while excluding orthogroups detected in only one genome. Applied with ANNEVO to 123 Aspergillus oryzae genomes, EukPan identified 11,245 core and 4,407 shared accessory orthogroups. The core-protein phylogeny broadly recovered the reported A-H classification, whereas accessory-genome analyses clearly separated the 33 group-A strains from the other 90 strains. Directional analysis identified 62 group-A-associated and 158 group-A-depleted orthogroups, with major facilitator superfamily (MFS) transporter and fungal Zn2Cys6 transcription-factor domains prominent in the depleted set. Among 93 orthogroups present in all non-A strains and absent from all group-A strains, 59 mapped to 10 segments of RIB40, the standard A. oryzae reference genome and a non-A (group-F) strain. EukPan therefore enables reproducible, coordinated core- and accessory-pangenome analysis from eukaryotic genome assemblies.

genomics↗