Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes

Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 to 6 MYA, but that are absent in the Hominidae. In fact, Hominidae show between four-and seven-fold lower rates of nucleotide change and feature turnover in both neutral and functional sequences suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. For example, recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli. This process resulted in thousands of novel, species-specific CTCF binding sites. Our results demonstrate that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.

genomics

Genome variation and conserved regulation identify genomic regions responsible for strain specific phenotypes in rat

The genomes of laboratory rat strains are characterised by a mosaic haplotype structure caused by their unique breeding history. These mosaic haplotypes have been recently mapped by extensive sequencing of key strains. Comparison of genomic variation between two closely related rat strains with different phenotypes has been proposed as an effective strategy for the discovery of candidate strain-specific regions involved in phenotypic differences.\n\nWe developed a method to prioritise strain-specific haplotypes by integrating genomic variation and genomic regulatory data predicted to be involved in specific phenotypes. To identify genomic regions associated with metabolic syndrome, a disorder of energy utilization and storage affecting several organ systems, we compared two Lyon rat strains, LH/Mav which is susceptible to MetS, and LL/Mav, which is susceptible to obesity as an intermediate MetS phenotype, with a third strain (LN/Mav) that is resistant to both MetS and obesity. Applying a novel metric, we ranked the identified strain-specific haplotypes using evolutionary conservation of the occupancy three liver-specific transcription factors (HNF4A, CEBPA, and FOXA1) in five rodents including rat.\n\nConsideration of regulatory information effectively identified regions with liver-associated genes and rat orthologues of human GWAS variants related to obesity and metabolic traits. We attempted to find possible causative variants and compared them with the candidate genes proposed by previous studies. In strain-specific regions with conserved regulation, we found a significant enrichment for published evidence to obesity--one of the metabolic symptoms shown by the Lyon strains--amongst the genes assigned to promoters with strain-specific variation.\n\nOur results show that the use of functional regulatory conservation is a potentially effective approach to select strain-specific genomic regions associated with phenotypic differences among Lyon rats and could be extended to other systems.

genomics

Leishmania naiffi and Leishmania guyanensis reference genomes highlight genome structure and gene content evolution in the Viannia subgenus

The unicellular protozoan parasite Leishmania causes the neglected tropical disease leishmaniasis, affecting 12 million people in 98 countries. In South America where the Viannia subgenus predominates, so far only L. (Viannia) braziliensis and L. (V.) panamensis have been sequenced, assembled and annotated as reference genomes. Addressing this deficit in molecular information can inform species typing, epidemiological monitoring and clinical treatment. Here, L. (V.) naiffi and L. (V.) guyanensis genomic DNA was sequenced to assemble these two genomes as draft references from short sequence reads. The methods used were tested using short sequence reads for L. braziliensis M2904 against its published reference as a comparison. This assembly and annotation pipeline identified 70 additional genes not annotated on the original M2904 reference. Phylogenetic and evolutionary comparisons of L. guyanensis and L. naiffi with ten other Viannia genomes revealed four traits common to all Viannia: aneuploidy, 22 orthologous groups of genes absent in other Leishmania subgenera, elevated TATE transposon copies, and a high NADH-dependent fumarate reductase gene copy number. Within the Viannia, there were limited structural changes in genome architecture specific to individual species: a 45 Kb amplification on chromosome 34 was present in all bar L. lainsoni, L. naiffi had a higher copy number of the virulence factor leishmanolysin, and laboratory isolate L. shawi M8408 had a possible minichromosome derived from the 3 end of chromosome 34. This combination of genome assembly, phylogenetics and comparative analysis across an extended panel of diverse Viannia has uncovered new insights into the origin and evolution of this subgenus and can help improve diagnostics for leishmaniasis surveillance.

genomics

The genomic false shuffle: epigenetic maintenance of topological domains in the rearranged gibbon genome

The relationship between evolutionary genome remodeling and the three-dimensional structure of the genome remain largely unexplored. Here we use the heavily rearranged gibbon genome to examine how evolutionary chromosomal rearrangements impact genome-wide chromatin interactions, topologically associating domains (TADs), and their epigenetic landscape. We use high-resolution maps of gibbon-human breaks of synteny (BOS), apply Hi-C in gibbon, measure an array of epigenetic features, and perform cross-species comparisons. We find that gibbon rearrangements occur at TAD boundaries, independent of the parameters used to identify TADs. This overlap is supported by a remarkable genetic and epigenetic similarity between BOS and TAD boundaries, namely presence of CpG islands and SINE elements, and enrichment in CTCF and H3K4me3 binding. Cross-species comparisons reveal that regions orthologous to BOS also correspond with boundaries of large (400-600kb) TADs in human and other mammalian species. The co-localization of rearrangement breakpoints and TAD boundaries may be due to higher chromatin fragility at these locations and/or increased selective pressure against rearrangements that disrupt TAD integrity. We also examine the small portion of BOS that did not overlap with TAD boundaries and gave rise to novel TADs in the gibbon genome. We postulate that these new TADs generally lack deleterious consequences. Lastly, we show that limited epigenetic homogenization occurs across breakpoints, irrespective of their time of occurrence in the gibbon lineage. Overall, our findings demonstrate remarkable conservation of chromatin interactions and epigenetic landscape in gibbons, in spite of extensive genomic shuffling.

genomics

Sequencing of the Venter/HuRef genome using various strategies for the benchmarking of genome analysis tools

We produced an extensive collection of deep re-sequencing datasets for the Venter/HuRef genome using the Illumina massively-parallel DNA sequencing platform. The original Venter genome sequence is a very-high quality phased assembly based on Sanger sequencing. Therefore, researchers developing novel computational tools for the analysis of human genome sequence variation for the dominant Illumina sequencing technology can test and hone their algorithms by making variant calls from these Venter/HuRef datasets and then immediately confirm the detected variants in the Sanger assembly, freeing them of the need for further experimental validation. This process also applies to implementing and benchmarking existing genome analysis pipelines. We prepared and sequenced 200 bp and 350 bp short-insert whole-genome sequencing libraries (sequenced to 100x and 40x genomic coverages respectively) as well as 2 kb, 5 kb, and 12 kb mate-pair libraries (49x, 122x, and 145x physical coverages respectively). Lastly, we produced a linked-read library (128x physical coverage) from which we also performed haplotype phasing.

genomics

The Cuon Enigma: Genome survey and comparative genomics of the endangered Dhole (Cuon alpinus)

The Asiatic wild dog is an endangered monophyletic canid restricted to Asia; facing threats from habitat fragmentation and other anthropogenic factors. Dholes have unique adaptations as compared to other wolf-like canids for large litter size (larger number of mammae) and hypercarnivory making it evolutionarily notable. Over evolutionary time, dhole and the subsequent divergent wild canids have lost coat patterns found in African wild dog. Here we report the first high coverage genome survey of Asiatic wild dog and mapped it with African wild dog, dingo and domestic dog to assess the structural variants. We generated a total of 124.8 Gb data from 416140921 raw read pairs and retained 398659457 reads with 52X coverage and mapped 99.16% of the clean reads to the three reference genomes. We identified ~13553269 SNVs, ~2858184 InDels, ~41000 SVs, ~1854109 SSRs and about 1000 CNVs. We compared the annotated genome of dingo and domestic dog with dhole genome sequence to understand the role of genes responsible in pelage pattern, dentition and mammary glands. Positively selected genes for these phenotypes were looked for SNP variants and top ranked genes for coat pattern, dentition and mammary glands were found to play a role in signalling and developmental pathways. Mitochondrial genome assembly predicted 35 genes, 11 CDS and 24 tRNA. This genome information will help in understanding the divergence of two monophlyletic canids, Cuon and Lycaon, and the evolutionary adaptations of dholes with respect to other canids.

genomics

The Sorghum bicolor reference genome: improved assembly and annotations, a transcriptome atlas, and signatures of genome organization

2Sorghum bicolor is a drought tolerant C4 grass used for production of grain, forage, sugar, and lignocellulosic biomass and a genetic model for C4 grasses due to its relatively small genome (~800 Mbp), diploid genetics, diverse germplasm, and colinearity with other C4 grass genomes. In this study, deep sequencing, genetic linkage analysis, and transcriptome data were used to produce and annotate a high quality reference genome sequence. Reference genome sequence order was improved, 29.6 Mbp of additional sequence was incorporated, the number of genes annotated increased 24% to 34,211, average gene length and N50 increased, and error frequency was reduced 10-fold to 1 per 100 kbp. Sub-telomeric repeats with characteristics of Tandem Repeats In Miniature (TRIM) elements were identified at the termini of most chromosomes. Nucleosome occupancy predictions identified nucleosomes positioned immediately downstream of transcription start sites and at different densities across chromosomes. Alignment of the reference genome sequence to 56 resequenced genomes from diverse sorghum genotypes identified ~7.4M SNPs and 1.8M indels. Large scale variant features in euchromatin were identified with periodicities of ~25 kbp. An RNA transcriptome atlas of gene expression was constructed from 47 samples derived from growing and developed tissues of the major plant organs (roots, leaves, stems, panicles, seed) collected during the juvenile, vegetative and reproductive phases. Analysis of the transcriptome data indicated that tissue type and protein kinase expression had large influences on transcriptional profile clustering. The updated assembly, annotation, and transcriptome data represent a resource for C4 grass research and crop improvement.

plant biology

Genome Wide Association Analyses Based On Broadly Different Specifications For Prior Distributions, Genomic Windows, And Estimation Methods

A popular strategy (EMMAX) for genome wide association (GWA) analysis fits all marker effects as classical random effects (i.e., Gaussian prior) by which association for the specific marker of interest is inferred by treating its effect as fixed. It seems more statistically coherent to specify all markers as sharing the same prior distribution, whether it is Gaussian, heavy-tailed (BayesA), or has variable selection specifications based on a mixture of, say, two Gaussian distributions (SSVS). Furthermore, all such GWA inference should be formally based on posterior probabilities or test statistics as we present here, rather than merely being based on point estimates. We compared these three broad categories of priors within a simulation study to investigate the effects of different degrees of skewness for quantitative trait loci (QTL) effects and numbers of QTL using 43,266 SNP marker genotypes from 922 Duroc-Pietrain F2 cross pigs. Genomic regions were based either on single SNP associations, on non-overlapping windows of various fixed sizes (0.5 to 3 Mb) or on adaptively determined windows that cluster the genome into blocks based on linkage disequilibrium (LD). We found that SSVS and BayesA lead to the best receiver operating curve properties in almost all cases. We also evaluated approximate marginal a posteriori (MAP) approaches to BayesA and SSVS as potential computationally feasible alternatives; however, MAP inferences were not promising, particularly due to their sensitivity to starting values. We determined that it is advantageous to use variable selection specifications based on adaptively constructed genomic window lengths for GWA studies.\n\nSUMMARYGenome wide association (GWA) analyses strategies have been improved by simultaneously fitting all marker effects when inferring upon any single marker effect, with the most popular distributional assumption being normality. Using data generated from 43,266 genotypes on 922 Duroc-Pietrain F2 cross pigs, we demonstrate that GWA studies could particularly benefit from more flexible heavy-tailed or variable selection distributional assumptions. Furthermore, these associations should not just be based on single markers or even genomic windows of markers of fixed physical distances (0.5 - 3.0 Mb) but based on adaptively determined genomic windows using linkage disequilibrium information.

genetics

Database-integrated genome screening (DIGS): exploring genomes heuristically using sequence similarity search tools and a relational database.

A significant fraction of most genomes is comprised of DNA sequences that have been incompletely investigated. This genomic dark matter contains a wealth of useful biological information that can be recovered by systematically screening genomes in silico using sequence similarity search tools. Specialized computational tools are required to implement these screens efficiently. Here, we describe the database-integrated genome-screening (DIGS) tool: a computational framework for performing these investigations. To demonstrate, we screen mammalian genomes for endogenous viral elements (EVEs) derived from the Filoviridae, Parvoviridae, Circoviridae and Bornaviridae families, identifying numerous novel elements in addition to those that have been described previously. The DIGS tool provides a simple, robust framework for implementing a broad range of heuristic, sequence analysis-based explorations of genomic diversity.\n\nAvailabilityhttp://giffordlabcvr.github.io/DIGS-tool/\n\nContactrobert.gifford@glasgow.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Systematic characterization of genome editing in primary T cells reveals proximal genomic insertions and enables machine learning prediction of CRISPR-Cas9 DNA repair outcomes

The Streptococcus pyogenes Cas9 (SpCas9) nuclease has become a ubiquitous genome editing tool due to its ability to target almost any location in DNA and create a double-stranded break1,2. After DNA cleavage, the break is fixed with endogenous DNA repair machinery, either by non-templated mechanisms (e.g. non-homologous end joining (NHEJ) or microhomology-mediated end joining (MMEJ)), or homology directed repair (HDR) using a complementary template sequence3,4. Previous work has shown that the distribution of repair outcomes within a cell population is non-random and dependent on the targeted sequence, and only recent efforts have begun to investigate this further5-11. However, no systematic work to date has been validated in primary human cells5,7. Here, we report DNA repair outcomes from 1,521 unique genomic locations edited with SpCas9 ribonucleoprotein complexes (RNPs) in primary human CD4+ T cells isolated from multiple healthy blood donors. We used targeted deep sequencing to measure the frequency distribution of repair outcomes for each guide RNA and discovered distinct features that drive individual repair outcomes after SpCas9 cleavage. Predictive features were combined into a new machine learning model, CRISPR Repair OUTcome (SPROUT), that predicts the length and probability of nucleotide insertions and deletions with R2 greater than 0.5. Surprisingly, we also observed large insertions at more than 90% of targeted loci, albeit at a low frequency. The inserted sequences aligned to diverse regions in the genome, and are enriched for sequences that are physically proximal to the break site due to chromatin interactions. This suggests a new mechanism where sequences from three-dimensionally neighboring regions of the genome can be inserted during DNA repair after Cas9-induced DNA breaks. Together, these findings provide powerful new predictive tools for Cas9-dependent genome editing and reveal new outcomes that can result from genome editing in primary T cells.

cell biology

Joint analysis of functional genomic data and genome-wide association studies of 18 human traits

Annotations of gene structures and regulatory elements can inform genome-wide association studies (GWAS). However, choosing the relevant annotations for interpreting an association study of a given trait remains challenging. We describe a statistical model that uses association statistics computed across the genome to identify classes of genomic element that are enriched or depleted for loci that influence a trait. The model naturally incorporates multiple types of annotations. We applied the model to GWAS of 18 human traits, including red blood cell traits, platelet traits, glucose levels, lipid levels, height, BMI, and Crohns disease. For each trait, we evaluated the relevance of 450 different genomic annotations, including protein-coding genes, enhancers, and DNase-I hypersensitive sites in over a hundred tissues and cell lines. We show that the fraction of phenotype-associated SNPs that influence protein sequence ranges from around 2% (for platelet volume) up to around 20% (for LDL cholesterol); that repressed chromatin is significantly depleted for SNPs associated with several traits; and that cell type-specific DNase-I hypersensitive sites are enriched for SNPs associated with several traits (for example, the spleen in platelet volume). Finally, by re-weighting each GWAS using information from functional genomics, we increase the number of loci with high-confidence associations by around 5%.

Genomics

Genome urbanization: Clusters of topologically co-regulated genes delineate functional compartments in the genome of S. cerevisiae

The eukaryotic genome evolves under the dual constraint of maintaining co-ordinated gene transcription and performing effective DNA replication and cell division, the coupling of which brings about inevitable DNA topological tension. DNA supercoiling is resolved and, in some cases, even harnessed by the genome through the function of DNA topoisomerases, as has been shown in the concurrent transcriptional activation and suppression of genes upon transient deactivation of topoisomerase II (topoII). By analyzing a genome wide run-on experiment upon thermal inactivation of topoII in S.cerevisiae. we were able to define 116 gene clusters of consistent response (either positive or negative) to topological stress. A comprehensive analysis of these topologically co-regulated gene clusters revealed pronounced preferences regarding their functional, regulatory and structural attributes. Genes that negatively respond to topological stress, are positioned in gene-dense pericentromeric regions, are more conserved and associated to essential functions, while up-regulated gene clusters are preferentially located in the gene-sparse nuclear periphery, associated with secondary functions and under complex regulatory control. We propose that evolves with a core of essential genes occupying a compact genomic \"old town\", whereas more recently acquired, condition-specific genes tend to be located in a more spacious \"suburban\" genomic periphery.

Genomics

Longitudinal genomic surveillance of Plasmodium falciparum malaria parasites reveals complex genomic architecture of emerging artemisinin resistance in western Thailand

BackgroundArtemisinin-based combination therapies are the first line of treatment for Plasmodium falciparum infections worldwide, but artemisinin resistance (ART-R) has risen rapidly in in Southeast Asia over the last decade. Mutations in kelch13 have been associated with artemisinin (ART) resistance in this region. To explore the power of longitudinal genomic surveillance to detect signals in kelch13 and other loci that contribute to ART or partner drug resistance, we retrospectively sequenced the genomes of 194 P. falciparum isolates from five sites in Northwest Thailand, bracketing the era in which there was a rapid increase in ART-R in this region (2001-2014).\n\nResultsWe evaluated statistical metrics for temporal change in the frequency of individual SNPs, assuming that SNPs associated with resistance should increase frequency over this period. After Kelch13-C580Y, the strongest temporal change was seen at a SNP in phosphatidylinositol 4-kinase (PI4K), situated in a pathway recently implicated in the ART-R mechanism. However, other loci exhibit temporal signatures nearly as strong, and warrant further investigation for involvement in ART-R evolution. Through genome-wide association analysis we also identified a variant in a kelch-domain-containing gene on chromosome 10 that may epistatically modulate ART-R.\n\nConclusionsThis analysis demonstrates the potential of a longitudinal genomic surveillance approach to detect resistance-associated loci and improve our mechanistic understanding of how resistance develops. Evidence for additional genomic regions outside of the kelch13 locus associated with ART-R parasites may yield new molecular markers for resistance surveillance and may retard the emergence or spread of ART-R in African parasite populations.

genomics

A new standard for crustacean genomes: the highly contiguous, annotated genome assembly of the clam shrimp Eulimnadia texana reveals HOX gene order and identifies the sex chromosome

Vernal pool clam shrimp (Eulimnadia texana) are a promising model system due to their ease of lab culture, short generation time, modest sized genome, a somewhat rare stable androdioecious sex determination system, and a requirement to reproduce via desiccated diapaused eggs. We generated a highly contiguous genome assembly using 46X of PacBio long read data and 216X of Illumina short reads, and annotated using Illumina RNAseq obtained from adult males or hermaphrodites. 85% of the 120Mb genome is contained in the largest 8 contigs, the smallest of which is 4.6Mb. The assembly contains 98% of transcripts predicted via RNAseq. This assembly is qualitatively different from scaffolded Illumina assemblies: it is produced from long reads that contain sequence data along their entire length, and is thus gap free. The contiguity of the assembly allows us to order the HOX genes within the genome, identifying two loci that contain HOX gene orthologs, and which approximately maintain the order observed in other arthropods. We identified a partial duplication of the Antennapedia gene adjacent to the few genes homologous to the Bithorax locus. Because the sex chromosome of an androdioecious species is of special interest, we used existing allozyme and microsatellite markers to identify the E. texana sex chromosome, and find that it comprises nearly half of the genome of this species. Linkage patterns indicate that recombination is extremely rare and perhaps absent in hermaphrodites, and as a result the location of the sex determining locus will be difficult to refine using recombination mapping.

genomics

Genome-wide enhancer - gene regulatory maps in two vertebrate genomes

The spatiotemporal expression of genes is controlled by enhancer sequences that bind transcription factors. Identifying the target genes of enhancers remains difficult because enhancers regulate gene expression over long genomic distances. To address this, we used an evolutionary approach to build two genome-wide maps of enhancer-gene associations in the human and zebrafish genomes. Enhancers were identified using sequence conservation, and linked to their predicted target genes using PEGASUS, a bioinformatics method that relies on evolutionary conservation of synteny. The analysis of these maps revealed that the number of enhancers linked to a gene correlate with its expression breadth. Comparison of both maps identified hundreds of vertebrate ancestral regulatory relationships from which we could determine that enhancer-gene distances scale with genome size despite strong positional conservation. The two maps represent a resource for further studies, including the prioritisation of sequence variants in whole genome sequence of patients affected by genetic diseases.

genomics

Linking the International Wheat Genome Sequencing Consortium bread wheat reference genome sequence to wheat genetic and phenomic data

The Wheat@URGI portal (https://wheat-urgi.versailles.inra.fr) has been developed to provide the international community of researchers and breeders with access to the bread wheat reference genome sequence produced by the International Wheat Genome Sequencing Consortium. Genome browsers, BLAST, and InterMine tools have been established for in depth exploration of the genome sequence together with additional linked datasets including physical maps, sequence variations, gene expression, and genetic and phenomic data from other international collaborative projects already stored in the GnpIS information system. The portal provides enhanced search and browser features that will facilitate the deployment of the latest genomics resources in wheat improvement.

bioinformatics

De Novo assembly of the goldfish (Carassius auratus) genome and the evolution of genes after whole genome duplication

For over a thousand years throughout Asia, the common goldfish (Carassius auratus) was raised for both food and as an ornamental pet. Selective breeding over more than 500 years has created a wide array of body and pigmentation variation particularly valued by ornamental fish enthusiasts. As a very close relative of the common carp (Cyprinus carpio), goldfish shares the recent genome duplication that occurred approximately 14-16 million years ago (mya) in their common ancestor. The combination of centuries of breeding and a wide array of interesting body morphologies is an exciting opportunity to link genotype to phenotype as well as understanding the dynamics of genome evolution and speciation. Here we generated a high-quality draft sequence of a \"Wakin\" goldfish using 71X PacBio long-reads. We identified 70,324 coding genes and more than 11,000 non-coding transcripts. We found that the two sub-genomes in goldfish retained extensive synteny and collinearity between goldfish and zebrafish. However, \"ohnologous\" genes were lost quickly after the carp whole-genome duplication, and the expression of 30% of the retained duplicated gene diverged significantly across seven tissues sampled. Loss of sequence identity and/or exons determined the divergence of the expression across all tissues, while loss of conserved, non-coding elements determined expression variance between different tissues. This draft assembly also provides an important resource for comparative genomics with the very commonly used zebrafish model (Danio rerio), and for understanding the underlying genetic causes of goldfish variants.

genomics

Genomic variants concurrently listed in a somatic and a germline mutation database have implications for disease-variant discovery and genomic privacy

BackgroundMutations arise in the human genome in two major settings: the germline and soma. These settings involve different inheritance patterns, chromatin structures, and environmental exposures, all of which might be predicted to differentially affect the distribution of substitutions found in these settings. Nonetheless, recent studies have found that somatic and germline mutation rates are similarly affected by endogenous mutational processes and epigenetic factors.\n\nResultsHere, we quantified the number of single nucleotide variants that co-occur between somatic and germline call-sets (cSNVs), compared this quantity with expectations, and explained noted departures. We found that three times as many variants are shared between the soma and germline than is expected by independence. We developed a new, general-purpose statistical framework to explain the observed excess of cSNVs in terms of the varying mutation rates of different kinds substitution types and of genomic regions. Using this metric, we find that more than 90% of this excess can be explained by our observation that the basic substitution types (such as N[C->T]G, C->A, etc.) have correlated mutation rates in the germline and soma. Matched-normal read depth analysis suggests that an appreciable fraction of this excess may also derive from germline contamination of somatic samples.\n\nConclusionOverall, our results highlight the commonalities in substitution patterns between the germline and soma. The universality of some aspects of human mutation rates offers insight into the potential molecular mechanisms of human mutation. The highlighted similarities between somatic and germline mutation rates also lay the groundwork for future studies that distinguish disease-causing variants from a genomic background informed by both somatic and germline variant data. Moreover, our results also indicate that the depth of matched normal sequencing necessary to ensure genomic privacy of donors of somatic samples may be higher than previously appreciated. Furthermore, the fact that we were able to explain such a high portion of recurrent variants using known determinants of mutation rates is evidence that the genomics community has already discovered the most important predictors of mutation rates for single nucleotide variants.

genomics