Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Cohesin facilitates zygotic genome activation in zebrafish

At zygotic genome activation (ZGA), changes in chromatin structure are associated with new transcription immediately following the maternal-to-zygotic transition (MZT). The nuclear architectural proteins, cohesin and CCCTC-binding factor (CTCF), contribute to chromatin structure and gene regulation. We show here that normal cohesin function is important for ZGA in zebrafish. Depletion of cohesin subunit Rad21 delays ZGA without affecting cell cycle progression. In contrast, CTCF depletion has little effect on ZGA whereas complete abrogation is lethal. Genome wide analysis of Rad21 binding reveals a change in distribution from pericentromeric satellite DNA, and few locations including the miR-430 locus (whose products are responsible for maternal transcript degradation), to genes, as embryos progress through the MZT. After MZT, a subset of Rad21 binding occurs at genes dysregulated upon Rad21 depletion and overlaps pioneer factor Pou5f3, which activates early expressed genes. Rad21 depletion disrupts the formation of nucleoli and RNA polymerase II foci, suggestive of global defects in chromosome architecture. We propose that Rad21/cohesin redistribution to active areas of the genome is key to the establishment of chromosome organization and the embryonic developmental program.\n\nAuthor SummaryDuring the first few hours of existence, early zygotic cellular events are regulated by maternally inherited molecules. From a defined timepoint, the zygotic genome gradually becomes active and is transcribed. How the zygotic genome is first held inactive before becoming rapidly activated is poorly understood. Both gene repression and activation mechanisms are involved, but one aspect that has not yet been investigated is how 3-dimensional chromosome structure influences genome activation. In this study, we used zebrafish embryos to model zygotic genome activation.\n\nThe multi-subunit protein complex, cohesin, and the DNA-binding protein CCCTC-binding factor (CTCF) both have well known and overlapping roles in 3-dimensional genome organization. We depleted cohesin subunit Rad21, or CTCF, to determine their effects on zygotic genome activation. Moderate Rad21 depletion delayed transition to zygotic gene expression, without disrupting the cell cycle. By contrast, moderate CTCF depletion had very little effect; however, strong depletion of CTCF was lethal. We surveyed genome-wide binding of Rad21 before and after the zygotic genome is activated, and determined what other chromatin factors and transcription factors coincide with Rad21 binding. Before genome activation, Rad21 was located at satellite DNA and a few noncoding genes, one of which (miR-430) is responsible for degrading maternal transcripts. Following genome activation, there was a mass relocation of Rad21 to genes, particularly active genes and those that are targets of transcriptional activators when the zygotic genome is switched on. Depletion of Rad21 also affected global chromosome structure.\n\nOur study shows that cohesin binding redistributes to active RNA Polymerase II genes at the onset of zygotic gene transcription. Furthermore, we suggest that cohesin contributes to dynamic changes in chromosome architecture that occur upon zygotic genome activation.

developmental biology

Long-read whole genome sequencing and comparative analysis of six strains of the human pathogen Orientia tsutsugamushi

BackgroundOrientia tsutsugamushi is a clinically important but neglected obligate intracellular bacterial pathogen of the Rickettsiaceae family that causes the potentially life-threatening human disease scrub typhus. In contrast to the genome reduction seen in many obligate intracellular bacteria, early genetic studies of Orientia have revealed one of the most repetitive bacterial genomes sequenced to date. The dramatic expansion of mobile elements has hampered efforts to generate complete genome sequences using short read sequencing methodologies, and consequently there have been few studies of the comparative genomics of this neglected species.\n\nResultsWe report new high-quality genomes of Orientia tsutsugamushi, generated using PacBio single molecule long read sequencing, for six strains: Karp, Kato, Gilliam, TA686, UT76 and UT176. In comparative genomics analyses of these strains together with existing reference genomes from Ikeda and Boryong strains, we identify a relatively small core genome of 657 genes, grouped into core gene islands and separated by repeat regions, and use the core genes to infer the first whole-genome phylogeny of Orientia.\n\nConclusionsComplete assemblies of multiple Orientia genomes verify initial suggestions that these are remarkable organisms. They have large genomes with widespread amplification of repeat elements and massive chromosomal rearrangements between strains. At the gene level, Orientia has a relatively small set of universally conserved genes, similar to other obligate intracellular bacteria, and the relative expansion in genome size can be accounted for by gene duplication and repeat amplification. Our study demonstrates the utility of long read sequencing to investigate complex bacterial genomes and characterise genomic variation.

genomics

A haplotype-resolved draft genome of the European sardine (Sardina pilchardus)

BackgroundThe European sardine (Sardina pilchardus Walbaum, 1792) has a high cultural and economic importance throughout its distribution. Monitoring studies of sardine populations report an alarming decrease in stocks due to overfishing and environmental change, which has resulted in historically low captures along the Iberian Atlantic coast. Consequently, there is an urgent need to better understand the causal factors of this continuing decrease in the sardine stock. Important biological and ecological features such as levels of population diversity, structure, and migratory patterns can be addressed with the development and use of genomics resources.\n\nFindingsThe sardine genome of a single female individual was sequenced using Illumina HiSeq X Ten 10X Genomics linked-reads generating 113.8 Gb of data. Three draft genomes were assembled: two haploid genomes with a total size of 935 Mbp (N50 103Kb) each, and a consensus genome with a total size of 950 Mbp (N50 97Kb). The genome completeness assessment captured 84% of Actinopterygii Benchmarking Universal Single-Copy Orthologs. To obtain a more complete analysis, the transcriptomes of eleven tissues were sequenced and used to aid the functional annotation of the genome, resulting in 40 777 genes predicted. Variant calling on nearly half of the haplotype genome resulted in the identification of more than 2.3 million phased SNPs with heterozygous loci.\n\nConclusionsA draft genome was obtained with the 10X Genomics linked-reads technology, despite a high level of sequence repeats and heterozygosity that are expected genome characteristics of a wild sardine. The reference sardine genome and respective variant data are a cornerstone resource of ongoing population genomics studies to be integrated into future sardine stock assessment modelling to better manage this valuable resource.

genomics

Genomic Relatedness Strengthens Genetic Connectedness Across Management Units

Genetic connectedness refers to a measure of genetic relatedness across management units (e.g., herds and flocks). With the presence of high genetic connectedness in management units, best linear unbiased prediction (BLUP) is known to provide reliable comparisons between genetic values. Genetic connectedness has been studied for pedigree-based BLUP; however, relatively little attention has been paid to using genomic information to measure connectedness. In this study, we assessed genome-based connectedness across management units by applying prediction error variance of difference (PEVD), coefficient of determination (CD), and prediction error correlation (r) to a combination of computer simulation and real data (mice and cattle). We found that genomic information (G) increased the estimate of connectedness among individuals from different management units compared to that based on pedigree (A). A disconnected design benefited the most. In both datasets, PEVD and CD statistics inferred increased connectedness across units when using G- rather than A-based relatedness suggesting stronger connectedness. With r once using allele frequencies equal to one-half or scaling G to values between 0 and 2, which is intrinsic to A, connectedness also increased with genomic information. However, PEVD occasionally increased, and r decreased when obtained using the alternative form of G, instead suggesting less connectedness. Such inconsistencies were not found with CD. We contend that genomic relatedness strengthens measures of genetic connectedness across units and has the potential to aid genomic evaluation of livestock species.\n\nThe problem of connectedness or disconnectedness is particularly important in genetic evaluation of managed populations such as domesticated livestock. When selecting among animals from different management units (e.g., herds and flocks), caution is needed; choosing one animal over others across management units may be associated with greater uncertainty than selection within management units. Such uncertainty is reduced if individuals from different management units are genetically linked or connected. In such a case, best linear unbiased prediction (BLUP) offers meaningful comparison of the breeding values across management units for genetic evaluation (e.g., Kuehn et al., 2007).\n\nStructures of breeding programs have a direct influence on levels of connectedness. Wide use of artificial insemination (AI) programs generally increases genetic connectedness across management units. For example, dairy cattle populations are considered highly connected due to dissemination of genetic material from a small number of highly selected sires. The situation may be different for species with less use of AI and more use of natural service mating such as for beef cattle or sheep populations. Under these scenarios, the magnitude of connectedness across management units is reduced and genetic links are largely confined within management units.\n\nPedigree-based genetic connectedness has been evaluated and applied in practice (e.g., Kuehn et al., 2009; Eikje and Lewis, 2015). However, there is a relative paucity of use of genomic information such as single nucletide polymorphisms (SNPs) to ascertain connectedness. It still remains elusive in what scenarios genomics can strengthen connectedness and how much gain can be expected relative to use of pedigree information alone. Connectedness statistics have been used to optimize selective genotyping and phenotyping in simulated livestock (Pszczola et al., 2012) and plant populations (Maenhout et al., 2010), and in real maize (Rincent et al., 2012; Isidro et al., 2015), and real rice data (Isidro et al., 2015). These studies concluded that the greater the connectedness between the reference and validation populations, the greater the predictive performance. However, 1) connectedness among different management units and 2) differences in connectedness measures between pedigree and genomic relatedness were not explored in those studies. For better understanding of genome-based connectedness, it is critical to examine how the presence of management units comes into play. For instance, genomic relatedness provides relationships between distant individuals that appear disconnected according to the pedigree information. In addition, it captures Mendelian sampling that is not present in pedigree relationships (Hill and Weir, 2011). Thus, genomic information is expected to strengthen measures of connectedness, which in turn refines comparisons of genetic values across different management units. The objective of this study was to assess measures of genetic connectedness across management units with use of genomic information. We leveraged the combination of real data and computer simulation to compare gains in measures of connectedness when moving from pedigree to genomic relationships. First, we studied a heterogenous mice dataset stratified by cage. Then we investigated approaches to measure connectedness using real cattle data coupled with simulated management units to have greater control over the degree of confounding between fixed management groups and genetic relationships.

genetics

Human demographic history has amplified the effects background selection across the genome

Natural populations often grow, shrink, and migrate over time. Demographic processes such as these can impact genome-wide levels of genetic diversity. In addition, genetic variation in functional regions of the genome can be altered by natural selection, which drives adaptive mutations to higher frequencies or purges deleterious ones. Such selective processes impact not only the sites directly under selection but also nearby neutral variation through genetic linkage through processes referred to as genetic hitch-hiking in the context of positive selection and background selection (BGS) in the context of purifying selection. While there is extensive literature examining the impact of selection at linked sites at demographic equilibrium, less is known about how non-equilibrium demographic processes impact the effects of hitchhiking and BGS. Utilizing a global sample of human whole-genome sequences from the Thousand Genomes Project and extensive simulations, we investigate how non-equilibrium demographic processes magnify and dampen the consequences of selection at linked sites across the human genome. When binning the genome by inferred strength of BGS, we observe that, compared to Africans, non-African populations have experienced larger proportional decreases in neutral genetic diversity in such regions. We replicate these findings in admixed populations by showing that non-African ancestral components of the genome have also been impacted more severely in these regions. We attribute these differences to the strong, sustained/recurrent population bottlenecks that non-Africans experienced as they migrated out of Africa and throughout the globe. Furthermore, we observe a strong correlation between FST and inferred strength of BGS, suggesting a stronger rate of genetic drift. Forward simulations of human demographic history with a model of BGS support these observations. Our results show that non-equilibrium demography significantly alters the consequences selection at linked sites and support the need for more work investigating the dynamic process of multiple evolutionary forces operating in concert.\n\nAuthor summaryPatterns of genetic diversity within a species are affected at broad and fine scales by population size changes (\"demography\") and natural selection. From both population genetics theory and observation of genomic sequence data, it is known that demography can alter genome-wide average neutral genetic diversity. Additionally, natural selection can affect neutral genetic diversity regionally across the genome via selection at linked sites. During this process, natural selection acting on adaptive or deleterious variants in the genome will also impact diversity at nearby neutral sites due to genetic linkage. However, less is well known about the dynamic changes to diversity that occur in regions impacted by selection at linked sites when a population undergoes a size change. We characterize these dynamic changes using thousands of human genomes and find that the population size changes experienced by humans have shaped the consequences of linked selection across the genome. In particular, population contractions, such as those experienced by non-Africans, have disproportionately decreased neutral diversity in regions of the genome inferred to be under strong background selection (i.e., selection at linked sites that is caused by natural selection acting on deleterious variants), resulting in large differences between African and non-African populations.

evolutionary biology

A planarian nidovirus expands the limits of RNA genome size

RNA viruses are the only known RNA-protein (RNP) entities capable of autonomous replication (albeit within a permissive environment). A 33.5-kb nidovirus has been considered close to the upper size limit for such entities; conversely, the minimal cellular DNA genome is ~200 kb. This large difference presents a daunting gap for the transition from primordial RNP to contemporary DNA-RNP-based life. Whether or not RNA viruses represent transitional steps on the road to DNA-based life, studies of larger RNA viruses advance our understanding of size constraints on RNP entities. For example, emergence of the largest previously known RNA genomes (20-34 kb in positive-stranded nidoviruses, including coronaviruses) is associated with a proofreading exoribonuclease encoded in the nidoviral open reading frame 1b (ORF1b). However, apparent constraints on the size of ORF1b, which encodes this and other key replicative enzymes, have been hypothesized to limit further expansion of viral RNA genomes. Here, we characterize a novel nidovirus (planarian secretory cell nidovirus; PSCNV) whose disproportionately large ORF1b-like region, and overall 41.1 kb genome, substantially extend the presumed limits on RNA genome size. This genome encodes a predicted 13,556-aa polyprotein in an unconventional single ORF, yet retains canonical nidoviral genome organization and expression, and key replicative domains. Our evolutionary analysis suggests that PSCNV diverged early from multi-ORF nidoviruses, and subsequently acquired additional genes, including those typical of large DNA viruses or hosts. PSCNVs greatly expanded genome, proteomic complexity, and unique features - impressive in themselves - attest to the likelihood of still-larger RNA genomes awaiting discovery.\n\nSignificance StatementRNA viruses are the only known RNA-protein (RNP) entities capable of autonomous replication. The upper genome size for such entities was assumed to be <35 kb; conversely, the minimal cellular DNA genome is ~200 kb. This large difference presents a daunting gap for the proposed evolution of contemporary DNA-RNP-based life from primordial RNP entities. Here, we describe a nidovirus from planarians, whose 41.1 kb genome is 23% larger than the largest known of RNA virus. The planarian secretory cell nidovirus has broken apparent constraints on the size of the genomic subregion that encodes core replication machinery, and has acquired genes not previously observed in RNA viruses. This virus challenges and advances our understanding of the limits to RNA genome size.

microbiology

The tempo of linked selection: emergence of a heterogeneous genomic landscape during a recent radiation of monkeyflowers

Speciation genomic studies aim to interpret patterns of genome-wide variation in light of the processes that give rise to new species. However, interpreting the genomic landscape of speciation is difficult, because many evolutionary processes can impact levels of variation. Facilitated by the first chromosome-level assembly for the group, we use whole-genome sequencing and simulations to shed light on the processes that have shaped the genomic landscape during a recent radiation of monkeyflowers. After inferring the phylogenetic relationships among the nine taxa in this radiation, we show that highly similar diversity ({pi}) and differentiation (FST) landscapes have emerged across the group. Variation in these landscapes was strongly predicted by the local density of functional elements and the recombination rate, suggesting that the landscapes have been shaped by widespread natural selection. Using the varying divergence times between pairs of taxa, we show that the correlations between FST and genome features arose almost immediately after a population split and have become stronger over time. Simulations of genomic landscape evolution suggest that background selection (i.e., selection against deleterious mutations) alone is too subtle to generate the observed patterns, but scenarios that involve positive selection and genetic incompatibilities are plausible alternative explanations. Finally, tests for introgression among these taxa reveal widespread evidence of heterogeneous selection against gene flow during this radiation. Thus, combined with existing evidence for adaptation in this system, we conclude that the correlation in FST among these taxa informs us about the genomic basis of adaptation and speciation in this system.\n\nAuthor summaryWhat can patterns of genome-wide variation tell us about the speciation process? The answer to this question depends upon our ability to infer the evolutionary processes underlying these patterns. This, however, is difficult, because many processes can leave similar footprints, but some have nothing to do with speciation per se. For example, many studies have found highly heterogeneous levels of genetic differentiation when comparing the genomes of emerging species. These patterns are often referred to as differentiation landscapes because they appear as a rugged topography of peaks and valleys as one scans across the genome. It has often been argued that selection against deleterious mutations, a process referred to as background selection, is primarily responsible for shaping differentiation landscapes early in speciation. If this hypothesis is correct, then it is unlikely that patterns of differentiation will reveal much about the genomic basis of speciation. However, using genome sequences from nine emerging species of monkeyflower coupled with simulations of genomic divergence, we show that it is unlikely that background selection is the primary architect of these landscapes. Rather, differentiation landscapes have probably been shaped by adaptation and gene flow, which are processes that are central to our understanding of speciation. Therefore, our work has important implications for our understanding of what patterns of differentiation can tell us about the genetic basis of adaptation and speciation.

evolutionary biology

Global Genome Nucleotide Excision Repair is Organised into Domains Promoting Efficient DNA Repair in Chromatin

The rates at which lesions are removed by DNA repair can vary widely throughout the genome with important implications for genomic stability. To study this, we measured the distribution of nucleotide excision repair (NER) rates for UV-induced lesions throughout the budding yeast genome. By plotting these repair rates in relation to genes and their associated flanking sequences, we reveal that in normal cells, genomic repair rates display a distinctive pattern, suggesting that DNA repair is highly organised within the genome. Furthermore, by comparing genome-wide DNA repair rates in wild-type cells, and cells defective in the global genome-NER (GG-NER) sub-pathway, we establish how this alters the distribution of NER rates throughout the genome. We also examined the genomic locations of GG-NER factor binding to chromatin before and after UV irradiation revealing that GG-NER is organised and initiated from specific genomic locations. At these sites, chromatin occupancy of the histone acetyl transferase Gcn5 is controlled by the GG-NER complex, which regulates histone H3 acetylation and chromatin structure, thereby promoting efficient DNA repair of UV-induced lesions. Chromatin remodeling during the GG-NER process is therefore organized into these genomic domains. Importantly, loss of Gcn5, significantly alters the genomic distribution of NER rates, a finding that has important implications for the effects of chromatin modifiers on the distribution of mutations that arise throughout the genome.

Genomics

Deep Sequencing of 10,000 Human Genomes

We report on the sequencing of 10,545 human genomes at 30-40x coverage with an emphasis on quality metrics and novel variant and sequence discovery. We find that 84% of an individual human genome can be sequenced confidently. This high confidence region includes 91.5% of exon sequence and 95.2% of known pathogenic variant positions. We present thedistribution of over 150 million single nucleotide variants in the coding and non-coding genome. Each newly sequenced genome contributes an average of 8,579 novel variants. In addition, each genome carries in average 0.7 Mb of sequence that is not found in the main build of the hg38 reference genome. The density of this catalog of variation allowed us to construct highresolution profiles that define genomic sites that are highly intolerant of genetic variation. These results indicate that the data generated by deep genome sequencing is of the quality necessary for clinical use.\n\nSignificance statementDeclining sequencing costs and new large-scale initiatives towards personalized medicine are driving a massive expansion in the number of human genomes being sequenced. Therefore, there is an urgent need to define quality standards for clinical use. This includes deep coverage and sequencing accuracy of an individuals genome, rather than aggregated coverage of data across a cohort or population. Our work represents the largest effort to date in sequencing human genomes at deep coverage with these new standards. This study identifies over 150 million human variants, a majority of them rare and unknown. Moreover, these data identify sites in the genome that are highly intolerant to variation - possibly essential for life or health. We conclude that high coverage genome sequencing provides accurate detail on human variation for discovery and for clinical applications.

Genomics

Umap and Bismap: quantifying genome and methylome mappability

MotivationShort-read sequencing enables assessment of genetic and biochemical traits of individual genomic regions, such as the location of genetic variation, protein binding, and chemical modifications. Every region in a genome assembly has a property called mappability which measures the extent to which it can be uniquely mapped by sequence reads. In regions of lower mappability, estimates of genomic and epigenomic characteristics from sequencing assays are less reliable. At best, sequencing assays will produce misleadingly low numbers of reads in these regions. At worst, these regions have increased susceptibility to spurious mapping from reads from other regions of the genome with sequencing errors or unexpected genetic variation. Bisulfite sequencing approaches used to identify DNA methylation exacerbate these problems by introducing large numbers of reads that map to multiple regions. While many tools consider mappability during the read mapping process, subsequent analysis often loses this information. Both to correct assumptions of uniformity in downstream analysis, and to identify regions where the analysis is less reliable, it is necessary to know the mappability of both ordinary and bisulfite-converted genomes.\n\nResultsWe introduce the Umap software for identifying uniquely mappable regions of any genome. Its Bismap extension identifies mappability of the bisulfite-converted genome. With a read length of 24 bp, 18.7% of the unmodified genome and 33.5% of the bisulfite-converted genome is not uniquely mappable. This complicates interpretation of functional genomics experiments using short-read sequencing, especially in regulatory regions. For example, 81% of human CpG islands overlap with regions that are not uniquely mappable. Similarly, in some ENCODE ChIP-seq datasets, up to 50% of peaks overlap with regions that are not uniquely mappable. We also explored differentially methylated regions from a case-control study and identified regions that were not uniquely mappable. In the widely used 450K methylation array, 4,230 probes are not uniquely mappable. Genome mappability is higher with longer sequencing reads, but most publicly available ChIP-seq and reduced representation bisulfite sequencing datasets have shorter reads. Therefore, uneven and low mappability remains a concern in a majority of existing data.\n\nAvailabilityA Umap and Bismap track hub for human genome assemblies GRCh37/hg19 and GRCh38/hg38, and mouse assemblies GRCm37/mm9 and GRCm38/mm10 is available at http://bismap.hoffmanlab.org for use with the UCSC and Ensembl genome browsers. We have deposited in Zenodo the current version of our software (https://doi.org/10.5281/zenodo.800648) and the mappability data used in this project (https://doi.org/10.5281/zenodo.800645). In addition, the software (https://bitbucket.org/hoffmanlab/umap) is freely available under the GNU General Public License, version 3 (GPLv3).\n\nContactmichael.hoffman@utoronto.ca

genomics

Integrating genomic resources for a threatened Caribbean coral (Orbicella faveolata) using a genetic linkage map developed from individual larval genotypes

Genomic methods are powerful tools for studying evolutionary responses to selection, but the application of these tools in non-model systems threatened by climate change has been limited by the availability of genomic resources in those systems. High-throughput DNA sequencing has enabled development of genome and transcriptome assemblies in non-model systems including reef-building corals, but the fragmented nature of early draft assemblies often obscures the relative positions of genes and genetic markers, and limits the functional interpretation of genomic studies in these systems. To address this limitation and improve genomic resources for the study of adaptation to ocean warming in corals, weve developed a genetic linkage map for the mountainous star coral, Orbicella faveolata. We analyzed genetic linkage among multilocus SNP genotypes to infer the relative positions of markers, transcripts, and genomic scaffolds in an integrated genomic map. To illustrate the utility of this resource, we tested for genetic associations with bleaching responses and fluorescence phenotypes, and estimated genome-wide patterns of population differentiation. Mapping the significant markers identified from these analyses in the integrated genomic resource identified hundreds of genes linked to significant markers, highlighting the utility of this resource for genomic studies of corals. The functional interpretations drawn from genomic studies are often limited by the availability of genomic resources linking genes to genetic markers. The resource developed in this study provides a framework for comparing genetic studies of O. faveolata across genotyping methods or references, and illustrates an approach for integrating genomic resources that may be broadly useful in other non-model systems.

genomics

The Genome of the Human Pathogen Candida albicans is Shaped by Mutation and Cryptic Sexual Recombination

The opportunistic fungal pathogen Candida albicans lacks a conventional sexual program and is thought to evolve, at least primarily, through the clonal acquisition of genetic changes. Here, we performed an analysis of heterozygous diploid genomes from 21 clinical isolates to determine the natural evolutionary processes acting on the C. albicans genome. Consistent with a model of inheritance by descent, most single nucleotide polymorphisms (SNPs) were shared between closely related strains. However, strain-specific SNPs and insertions/deletions (indels) were distributed non-randomly across the genome. For example, base substitution rates were higher in the immediate vicinity of indels, and heterozygous regions of the genome contained significantly more strain-specific polymorphisms than homozygous regions. Loss of heterozygosity (LOH) events also contributed substantially to genotypic variation, with most long-tract LOH events extending to the ends of the chromosomes suggestive of repair via break-induced replication. Importantly, some isolates contained highly mosaic genomes and failed to cluster closely with other isolates within their assigned clades. Mosaicism is consistent with strains having experienced inter-clade recombination during their evolutionary history and a detailed examination of nuclear and mitochondrial genomes revealed striking examples of recombination. Together, our analyses reveal that both (para)sexual recombination and mitotic mutational processes drive evolution of this important pathogen in nature. To further facilitate the study of genome differences we also introduce an online platform, SNPMap, to examine SNP patterns in sequenced C. albicans genomes.\n\nAUTHOR SUMMARYMutations introduce variation into the genome upon which selection can act. Defining the nature of these changes is critical for determining species evolution, as well as for understanding the genetic changes driving important cellular processes such as carcinogenesis. The fungus Candida albicans is a heterozygous diploid species that is both a frequent commensal organism and a prevalent opportunistic pathogen. Prevailing theory is that C. albicans evolves primarily through the gradual build-up of mutations, and a pressing question is whether sexual or parasexual processes also operate within natural populations. Here, we determine the evolutionary patterns of genetic change that have accompanied species evolution in nature by examining genomic differences between clinical isolates. We establish that the C. albicans genome evolves by a combination of base-substitution mutations, insertions/deletion events, and both short-tract and long-tract loss of heterozygosity (LOH) events. These mutations are unevenly distributed across the genome, and reveal that non-coding regions and heterozygous regions are evolving more quickly than coding regions and homozygous regions, respectively. Furthermore, we provide evidence that genetic exchange has occurred between isolates, establishing that sexual or parasexual processes have transpired in C. albicans populations and contribute to the diversity of both nuclear and mitochondrial genomes.

genomics

Multi-species mosaicism of evolutionary origins of genomic loci harboring 59,732 human-specific regulatory sequences reflects a complex continuous speciation process of the human lineage

Nearly sixty thousand genomic regions harboring various types of candidate human-specific regulatory sequences (HSRS) have been identified using high-resolution sequencing technologies and methodologically diverse comparative analyses of human and non-human primates reference genomes. Here, the systematic analysis of evolutionary origins of 59,732 genomic loci harboring candidate HSRS has been performed to identify genomic sequences that were either inherited from extinct common ancestors (ECAs) or created de novo in human genomes after the split of human and chimpanzee lineages. Present analyses revealed thousands of HSRS that appear inherited from ECAs yet bypassed genomes of our closest evolutionary relatives, Chimpanzee and Bonobo, presumably due to the incomplete lineage sorting and/or species-specific loss or regulatory DNA. The bypassing pattern is particularly prominent for HSRS putatively associated with development and functions of human brain. Significant fractions of retrotransposon-derived loci that are transcriptionally-active in human dorsolateral prefrontal cortex are highly conserved in genomes of Gorilla, Orangutan, Gibbon, and Rhesus (1,688; 1,371; 1,148; and 1,045 loci, respectively), yet they are absent in genomes of both Chimpanzee and Bonobo. These observations were independently corroborated by common evolutionary patterns of 248 insertions sites of African ape-specific retrovirus PtERV1 (45.9%; p = 1.03E-44) intersecting genomic regions harboring 442 HSRS, which are enriched for HSRS that have been associated with human-specific (HS) changes of gene expression in cerebral organoid models of brain development. A prominent majority of genomic regions harboring HS mutations associated with HS gene expression changes during brain development is highly conserved in Chimpanzee, Bonobo, and Gorilla genomes. Among nonhuman primates (NHP), most significant fractions of candidate HSRS associated with HS gene expression changes in both excitatory neurons (347 loci; 67%) and radial glia (683 loci; 72%) are highly conserved in the Gorilla genome. Present analyses revealed that Modern Humans captured unique combinations of regulatory sequences, divergent subsets of which are highly conserved in distinct species of six NHP separated by 30 million years of evolution. Concurrently, this unique-to-human mosaic of genomic regulatory patterns inherited from ECAs was supplemented with 12,486 created de novo HSRS. Evidence of multispecies evolutionary origins of HSRS support the model of complex continuous speciation process during evolution of Great Apes that is not likely to occur as an instantaneous event.

genomics

Specific Virus-Host Genome Interactions Revealed By Tethered Chromosome Conformation Capture

Viruses have evolved a variety of mechanisms to interact with host cells for their adaptive benefits, including subverting host immune responses and hijacking host DNA replication/transcription machineries [1-3]. Although interactions between viral and host proteins have been studied extensively, little is known about how the vial genome may interact with the host genome and how such interactions could affect the activities of both the virus and the host cell. Since the three-dimensional organization of a genome can have significant impact on genomic activities such as transcription and replication, we hypothesize that such structure-based regulation of genomic functions also applies to viral genomes depending on their association with host genomic regions and their spatial locations inside the nucleus. Here, we used Tethered Chromosome Conformation Capture (TCC) to investigate viral-host genome interactions between the adenovirus and human lung fibroblast cells. We found viral-host genome interactions were enriched in certain active chromatin regions and chromatin domains marked by H3K27me3. The contacts by viral DNA seems to impact the structure and function of the host genome, leading to remodeling of the fibroblast epigenome. Our study represents the first comprehensive analysis of viral-host interactions at the genome structure level, revealing unexpectedly specific virus-host genome interactions. The non-random nature of such interactions indicates a deliberate but poorly understood mechanism for targeting of host DNA by foreign genomes.

microbiology

Unexpected Properties of Short Genomic Tandem Repeats

Length polymorphisms in genomic short tandem repeats have been implicated in a variety of diseases, most notably human neurodegenerative disorders. Expansions of tandem repeats are also associated with genomic instability in cancer. Our previous study of length-3 tandem repeats uncovered a surprising pattern in the length distribution of certain such repeats in the non-coding regions of the human reference genome: a bias towards repeats of length 3n - 1, (n > 3). That is, the observed frequency of repeats of this length in the human genome is higher than expected by chance based on the frequency of shorter repeats.\n\nWe have hypothesized that this pattern may be a general property of genomic DNA. If true, this could have implications with regard to the dynamics of repeat expansion generally. To test this hypothesis, we have analyzed the genomic sequences of a broad range of eukaryotic organisms as well as several complete human genomes and obtained a number of thought provoking results. We establish that this unexpected elevation in frequency of 3n - 1 long repeats is statistically significant. We also expanded this analysis to different classes of genomic regions and tandem repeats of length four and five. The specific pattern was found in 13 of the 20 organisms analyzed, including all chordate and insect genomes tested. The bias pattern, however, was not confined to a single branch of the evolutionary tree. For some genomes, such as Drosophila melanogaster, the repeat bias surprisingly was also identified in exons. The pattern is present in both small and large genomes. A similar pattern was also found in tetranucleotide and pentanucleotide repeats in the human genome. Another surprising property was identified for the flanking GC content for triplet repeats of length 3n. These findings indicate a puzzling new genomic phenomenon with possible evolutionary and disease-related implications.

bioinformatics

Ecological selection for small microbial genomes along a temperate-to-thermal soil gradient

Small bacterial and archaeal genomes provide insights into the minimal requirements for life1 and seem to be widespread on the microbial phylogenetic tree2. We know that evolutionary processes, mainly selection and drift, can result in microbial genome reduction 3,4. However, we do not know the precise environmental pressures that constrain genome size in free-living microorganisms. A study including isolates 5 has shown that bacteria with high optimum growth temperatures, including thermophiles, often have small genomes 6. It is unclear how well this relationship may extend generally to microorganisms in nature 7,8, and in particular to those microbes inhabiting complex and highly variable environments like soil 3,6,9. To understand the genomic traits of thermally-adapted microorganisms, here we investigated bacterial and archaeal metagenomes from a 45{degrees}C gradient of temperate-to-thermal soils overlying the ongoing Centralia, Pennsylvania (USA) coal seam fire. There was a strong relationship between average genome size and temperature: hot soils had small genomes relative to ambient soils (Pearsons r = -0.910, p < 0.001). There was also an inverse relationship between soil temperature and cell size (Pearsons r = -0.65, p = 0.021), providing evidence that cell and genome size in the wild are together constrained by temperature. Notably, hot soils had different community structures than ambient soils, implicating ecological selection for thermo-tolerant cells that had small genomes, rather than contemporary genome streamlining within the local populations. Hot soils notably lacked genes for described two-component regulatory systems and antimicrobial production and resistance. Our work provides field evidence for the inverse relationship between microbial genome size and temperature requirements in a diverse, free-living community over a wide range of temperatures that support microbial life. Our findings demonstrate that ecological selection for thermophiles and thermo-tolerant microorganisms can result in smaller average genome sizes in situ, possibly because they have small genomes reminiscent of a more ancestral state.

microbiology

Thermosipho spp. immune system differences affect variation in genome size and geographical distributions

Thermosipho species inhabit thermal environments such as marine hydrothermal vents, petroleum reservoirs and terrestrial hot springs. A 16S rRNA phylogeny of available Thermosipho spp. sequences suggested habitat specialists adapted to living in hydrothermal vents only, and habitat generalists inhabiting oil reservoirs, hydrothermal vents and hotsprings. Comparative genomics of 15 Thermosipho genomes separated them into three distinct species with different habitat distributions: the widely distributed T. africanus and the more specialized, T. melanesiensis and T. affectus. Moreover, the species can be differentiated on the basis of genome size, genome content and immune system composition. For instance, the T. africanus genomes are largest and contained the most carbohydrate metabolism genes, which could explain why these isolates were obtained from ecologically more divergent habitats. Nonetheless, all the Thermosipho genomes, like other Thermotogae genomes, show evidence of genome streamlining. Genome size differences between the species could further be correlated to differences in defense capacities against foreign DNA, which influence recombination via HGT. The smallest genomes are found in T. affectus that contain both CRISPR-cas Type I and III systems, but no RM system genes. We suggest that this has caused these genomes to be almost devoid of mobile elements, contrasting the two other species genomes that contain a higher abundance of mobile elements combined with different immune system configurations. Taken together, the comparative genomic analyses of Thermosipho spp. revealed genetic variation allowing habitat differentiation within the genus as well as differentiation with respect to invading mobile DNA.

microbiology

Genomic differentiation is initiated without physical linkage among targets of divergent selection in Fall armyworms

The process of speciation involves whole genome differentiation by overcoming gene flow between diverging populations. We have ample knowledge which evolutionary forces may cause genomic differentiation, and several speciation models have been proposed to explain the transition from genetic to genomic differentiation. However, it is still unclear what are critical conditions enabling genomic differentiation in nature. The Fall armyworm, Spodoptera frugiperda, is observed as two sympatric strains that have different host-plant ranges, suggesting the possibility of ecological divergent selection. In our previous study, we observed that these two strains show genetic differentiation across the whole genome with an unprecedentedly low extent, suggesting the possibility that whole genome sequences started to be differentiated between the strains. In this study, we analyzed whole genome sequences from these two strains from Mississippi to identify critical evolutionary factors for genomic differentiation. The genomic Fst is low (0.017) while 91.3% of 10kb windows have Fst greater than 0, suggesting genome-wide differentiation with a low extent. We identified nearly 400 outliers of genetic differentiation between strains, and found that physical linkage among these outliers is not a primary cause of genomic differentiation. Fst is not significantly correlated with gene density, a proxy for the strength of selection, suggesting that a genomic reduction in migration rate dominates the extent of local genetic differentiation. Our analyses reveal that divergent selection alone is sufficient to generate genomic differentiation, and any following diversifying factors may increase the level of genetic differentiation between diverging strains in the process of speciation.

evolutionary biology