Search bioRxiv⌕ Search

Biology subjects

Schall, P. Z.

Publications and source records attributed to Schall, P. Z..

3 recordsLinked to original sources

Integrative genotyping and analysis of canine structural variation using long-read and short-read data

Structural variation makes an important contribution to canine evolution and phenotypic differences. Although recent advances in long-read sequencing have enabled the generation of multiple canine genome assemblies, most prior analyses of structural variation have relied on short read sequencing. To offer a more complete assessment of structural variation in canines, we performed an integrative analysis of structural variants present in 12 canine samples with available long-read and short-read sequencing data along with genome assemblies. Use of long-reads permits the discovery of heterozygous variation that is absent in existing haploid assembly representations while offering a marked increase in the ability to identify insertion variants relative to short-read approaches. Examination of the size spectrum of structural variants shows that dimorphic LINE-1 and SINE variants account for over 45% of all deletions and identified 1,410 LINE-1s with intact open reading frames that show presence-absence dimorphism. Using a graph-based approach, we genotype newly discovered structural variants in an existing collection of 1,879 resequenced dogs and wolves, generating a variant catalog containing a 56.5% increase in the number of deletions and 705% increase in the number of insertions previously found in the analyzed samples. Examination of allele frequencies across admixture components present across breed clades identified 283 structural variants evolving with a signature of selection. Significance statementExisting studies of structural variation have focused on genomes sequenced with short-read technology. In this study, we systematically combine short-read and long-read data sets to identify and genotype structural variation across canines, resulting in an expanded catalog of structural variants. Analysis of allele frequencies identified structural variants in genic regions that may have been evolving under selection across breed clades.

genomics↗

Characterization of nuclear mitochondrial insertions in canine genome assemblies

BackgroundThe presence of mitochondrial sequences in the nuclear genome (Numts) confounds analyses of mitochondrial sequence variation and is a potential source of false positives in disease studies. To improve the analysis of mitochondrial variation in canines, we completed a systematic assessment of Numt content across genome assemblies, canine populations and the carnivore lineage. ResultsCentering our analysis on the UU_Cfam_GSD_1.0/canFam4/Mischka assembly, a commonly used reference in dog genetic variation studies, we find a total of 321 Numts, located throughout the nuclear genome and encompassing the entire sequence of the mitochondria. Comparison to 14 canine genome assemblies identified 63 Numts with presence-absence dimorphism among dogs, wolves, and a coyote. Further, a subset of Numts were maintained across carnivore evolutionary time (arctic fox, polar bear, cat), with 8 sequences likely more than 10 million years old, and shared with the domestic cat. On a population level, using structural variant data from the Dog10K Consortium for 1,879 dogs and wolves, we identified 11 Numts that are absent in at least one sample as well as 53 Numts that are absent from the Mischka assembly. ConclusionsWe highlight scenarios where the presence of Numts is a potentially confounding factor and provide an annotation of these sequences in canine genome assemblies. This resource will aid the identification and interpretation of polymorphisms in both somatic and germline mitochondrial studies in canines.

genomics↗

A map of canine sequence variation relative to a Greenland wolf outgroup

For over 15 years, canine genetics research relied on a reference assembly from a Boxer breed dog named Tasha (i.e., canFam3.1). Recent advances in long-read sequencing and genome assembly have led to the development of numerous high-quality assemblies from diverse canines. These assemblies represent notable improvements in completeness, contiguity, and the representation of gene promoters and gene models. Although genome graph and pan-genome approaches have promise, most genetic analyses in canines rely upon the mapping of Illumina sequencing reads to a single reference. The Dog10K consortium, and others, have generated deep catalogs of genetic variation through an alignment of Illumina sequencing reads to a reference genome obtained from a German Shepherd Dog named Mischka (i.e., canFam4, UU_Cfam_GSD_1.0). However, alignment to a breed-derived genome may introduce bias in genotype calling across samples. Since the use of an outgroup reference genome may remove this effect, we have reprocessed 1,929 samples analyzed by the Dog10K consortium using a Greenland wolf (mCanLor1.2) as the reference. We efficiently performed remapping and variant calling using a GPU-implementation of common analysis tools. The resulting call set removes the variability in genetic differences seen across samples and breed relationships revealed by principal component analysis are not affected by the choice of reference genome. Using this sequence data, we inferred the history of population sizes and found that village dog populations experienced a 9-13 fold reduction in historic effective population size relative to wolves.

genomics↗