Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The Drosophila Genome Nexus: a population genomic resource of 605 Drosophila melanogaster genomes, including 197 genomes from a single ancestral range population

Hundreds of wild-derived D. melanogaster genomes have been published, but rigorous comparisons across data sets are precluded by differences in alignment methodology. The most common approach to reference-based genome assembly is a single round of alignment followed by quality filtering and variant detection. We evaluated variations and extensions of this approach, and settled on an assembly strategy that utilizes two alignment programs and incorporates both SNPs and short indels to construct an updated reference for a second round of mapping prior to final variant detection. Utilizing this approach, we reassembled published D. melanogaster population genomic data sets (previous DPGP releases and the DGRP freeze 2.0), and added unpublished genomes from several sub-Saharan populations. Most notably, we present aligned data from phase 3 of the Drosophila Population Genomics Project (DPGP3), which provides 197 genomes from a single ancestral range population of D. melanogaster (from Zambia). The large sample size, high genetic diversity, and potentially simpler demographic history of the DPGP3 sample will make this a highly valuable resource for fundamental population genetic research. The complete set of assemblies described here, termed the Drosophila Genome Nexus, presently comprises 605 consistently aligned genomes, and is publicly available in multiple formats with supporting documentation and bioinformatic tools. This resource will greatly facilitate population genomic analysis in this model species by reducing the methodological differences between data sets.

Genomics

Genome-wide Comparative Analysis Reveals Possible Common Ancestors of NBS Domain Containing Genes in Hybrid Citrus sinensis Genome and Original Citrus clementina Genome

Background Recently available whole genome sequences of three citrus species: one Citrus clementina and two Citrus sinensis genomes have made it possible to understand the features of candidate disease resistance genes with nucleotide-binding sites (NBS) domain in Citrus and how NBS genes differ between hybrid and original Citrus species. Result We identified and re-annotated NBS genes from three citrus genomes and found similar numbers of NBS genes in those citrus genomes. Phylogenetic analysis of all citrus NBS genes across three genomes showed that there are three approximately evenly numbered groups: one group contains the Toll-Interleukin receptor (TIR) domain and two different groups that contain the Coiled Coil (CC) domain. Motif analysis confirmed that the two groups of CC-containing NBS genes are from different evolutionary origins. We partitioned NBS genes into clades using NBS domain sequence distances and found most clades include NBS genes from all three citrus genomes. This suggests that NBS genes in three citrus genomes may come from shared ancestral origins. We also mapped the re-sequenced reads of three pomelo and three Mandarin orange genomes onto the Citrus sinensis genome. We found that most NBS genes of the hybrid C. sinensis genome have corresponding homologous genes in both pomelo and mandarin genome. The homologous NBS genes in pomelo and mandarin may explain why the NBS genes in their hybrid Citrus sinensis are similar to those in Citrus clementina in this study. Furthermore, sequence variation amongst citrus NBS genes were shaped by multiple independent and shared accelerated mutation accumulation events among different groups of NBS genes and in different citrus genomes. Conclusion Our comparative analyses yield valuable insight into the understanding of the structure, evolution and organization of NBS genes in Citrus genomes. There are significantly more NBS genes in Citrus genomes compared to other plant species. NBS genes in hybrid C. sinensis genomes are very similar to those in progenitor C. clementina genome and they may be derived from possible common ancestral gene copies. Furthermore, our comprehensive analysis showed that there are three groups of plant NBS genes while CC-containing NBS genes can be divided into two groups.

Genomics

Whole genome sequencing and assembly of a Caenorhabditis elegans genome with complex genomic rearrangements using the MinION sequencing device

Advances in 3rd generation sequencing have opened new possibilities for benchtop whole genome sequencing. The MinION is a portable device that uses nanopore technology and can sequence long DNA molecules. MinION long reads are well suited for sequencing and de novo assembly of complex genomes with large repetitive elements. Long reads also facilitate the identification of complex genomic rearrangements such as those observed in tumor genomes. To assess the feasibility of the de novo assembly of large complex genomes using both MinION and Illumina platforms, we sequenced the genome of a Caenorhabditis elegans strain that contains a complex acetaldehyde-induced rearrangement and a biolistic bombardment-mediated insertion of a GFP containing plasmid. Using [~]5.8 gigabases of MinION sequence data, we were able to assemble a C. elegans genome containing 145 contigs (N50 contig length = 1.22 Mb) that covered >99% of the 100,286,401 bp reference genome. In contrast, using [~]8.04 gigabases of Illumina sequence data, we were able to assemble a C. elegans genome in 38,645 contigs (N50 contig length = [~]26 kb) containing 117 Mb. From the MinION genome assembly we identified the complex structures of both the acetaldehyde-induced mutation and the biolistic-mediated insertion. To date, this is the largest genome to be assembled exclusively from MinION data and is the first demonstration that the long reads of MinION sequencing can be used for whole genome assembly of large (100 Mb) genomes and the elucidation of complex genomic rearrangements.

genomics

Beyond The Linear Genome: Comprehensive Determination Of The Endogenous Circular Elements In C. elegans And Human Genomes Via An Unbiased Genomic-Biophysical Method

Investigations aimed at defining the 3-D configuration of eukaryotic chromosomes have consistently encountered an endogenous population of chromosome-derived circular genomic DNA, referred to as extrachromosomal circular DNA (eccDNA). While the production, distribution, and activities of eccDNAs remain understudied, eccDNA formation from specific regions of the linear genome has profound consequences on the regulatory and coding capabilities for these regions. High-throughput sequencing has only recently made extensive genomic mapping of eccDNA sequences possible and had yet to be applied using a rigorous approach that distinguishes ascertainment bias from true enrichment. Here, we define eccDNA distribution, utilizing a set of unbiased topology-dependent approaches for enrichment and characterization. We use parallel biophysical, enzymatic, and informatic approaches to obtain a comprehensive profiling of eccDNA in C. elegans and in three human cell types, where eccDNAs were previously uncharacterized. We also provide quantitative analysis of the eccDNA loci at both unique and repetitive regions. Our studies converge on and support a consistent picture in which endogenous genomic DNA circles are present in normal physiological DNA metabolism, and in which the circles come from both coding and noncoding genomic regions. Prominent among the coding regions generating DNA circles are several genes known to produce a diversity of protein isoforms, with mucin proteins and titin as specific examples.

genomics

Genomes from uncultivated prokaryotes: a comparison of metagenome-assembled and single-amplified genomes

BackgroundProkaryotes dominate the biosphere and regulate biogeochemical processes essential to all life. Yet, our knowledge about their biology is for the most part limited to the minority that has been successfully cultured. Molecular techniques now allow for obtaining genome sequences of uncultivated prokaryotic taxa, facilitating in-depth analyses that may ultimately improve our understanding of these key organisms.\n\nResultsWe compared results from two culture-independent strategies for recovering bacterial genomes: single-amplified genomes and metagenome-assembled genomes. Single-amplified genomes were obtained from samples collected at an offshore station in the Baltic Sea Proper and compared to previously obtained metagenome-assembled genomes from a time series at the same station. Among 16 single-amplified genomes analyzed, seven were found to match metagenome-assembled genomes, affiliated with a diverse set of taxa. Notably, genome pairs between the two approaches were nearly identical (>98.7% identity) across overlapping regions (30-80% of each genome). Within matching pairs, the single-amplified genomes were consistently smaller and less complete, whereas the genetic functional profiles were maintained. For the metagenome-assembled genomes, only on average 3.6% of the bases were estimated to be missing from the genomes due to wrongly binned contigs; the metagenome assembly was found to cause incompleteness to a higher degree than the binning procedure.\n\nConclusionsThe strong agreement between the single-amplified and metagenome-assembled genomes emphasizes that both methods generate accurate genome information from uncultivated bacteria. Importantly, this implies that the research questions and the available resources are allowed to determine the selection of genomics approach for microbiome studies.

bioinformatics

Genome-wide homology analysis reveals new insights into the origin of the wheat B genome

Wheat is a typical allopolyploid with three homoeologous subgenomes (A, B, and D). The ancestors of the subgenomes A and D had been identified, but not for the subgenome B. The goatgrass Aegilops speltoides (genome SS) has been controversially considered a candidate for the ancestor of the wheat B genome. However, the relationship of the Ae. speltoides S genome with the wheat B genome remains largely obscure, which has puzzled the wheat research community for nearly a century. In the present study, the genome-wide homology analysis identified perceptible homology between wheat chromosome 1B and Ae. speltoides chromosome 1S, but not between other chromosomes in the B and S genomes. An Ae. speltoides-originated segment spanning a genomic region of approximately 10.46 Mb was identified on the long arm of wheat chromosome 1B (1BL). The Ae. speltoides-originated segment on 1BL was found to co-evolve with the rest of the B genome in wheat species. Thereby, we conclude that Ae. speltoides had been involved in the origin of the wheat B genome, but should not be considered an exclusive ancestor of this genome. The wheat B genome might have a polyphyletic origin with multiple ancestors involved, including Ae. speltoides. These novel findings provide significant insights into the origin and evolution of the wheat B genome, and will facilitate polyploid genome studies in wheat and other plants as well.

genetics

The whale shark genome reveals how genomic and physiological properties scale with body size

The endangered whale shark (Rhincodon typus) is the largest fish on Earth and is a long-lived member of the ancient Elasmobranchii clade. To characterize the relationship between genome features and biological traits, we sequenced and assembled the genome of the whale shark and compared its genomic and physiological features to those of 81 animals and yeast. We examined scaling relationships between body size, temperature, metabolic rates, and genomic features and found both general correlations across the animal kingdom and features specific to the whale shark genome. Among animals, increased lifespan is positively correlated to body size and metabolic rate. Several genomic features also significantly correlated with body size, including intron and gene length. Our large-scale comparative genomic analysis uncovered general features of metazoan genome architecture: GC content and codon adaptation index are negatively correlated, and neural connectivity genes are longer than average genes in most genomes. Focusing on the whale shark genome, we identified multiple features that significantly correlate with lifespan. Among these were very long gene length, due to large introns highly enriched in repetitive elements such as CR1-like LINEs, and considerably longer neural genes of several types, including connectivity, activity, and neurodegeneration genes. The whale sharks genome had an expansion of gene families related to fatty acid metabolism and neurogenesis, with the slowest evolutionary rate observed in vertebrates to date. Our comparative genomics approach uncovered multiple genetic features associated with body size, metabolic rate, and lifespan, and showed that the whale shark is a promising model for studies of neural architecture and lifespan.

genomics

The genomic landscape at a late stage of stickleback speciation: high genomic divergence interspersed by small localized regions of introgression

Speciation is a continuous process and analysis of species pairs at different stages of divergence provides insight into how it unfolds. Genomic studies on young species pairs have often revealed peaks of divergence and heterogeneous genomic differentiation. Yet it remains unclear how localised peaks of differentiation progress to genome-wide divergence during the later stages of speciation with gene flow. Spanning the speciation continuum, stickleback species pairs are ideal for investigating how genomic divergence builds up during speciation. However, attention has largely focused on young postglacial species pairs, with little known of the genomic signatures of divergence and introgression in older systems. The Japanese stickleback species pair, composed of the Pacific Ocean three-spined stickleback (Gasterosteus aculeatus) and the Japan Sea stickleback (G. nipponicus), which co-occur in the Japanese islands, is at a late stage of speciation. Divergence likely started well before the end of the last glacial period and crosses between Japan Sea females and Pacific Ocean males result in hybrid male sterility. Here we use coalescent analyses and Approximate Bayesian computation to show that the two species split approximately 0.68-1 million years ago but that they have continued to hybridise at a low rate throughout divergence. Population genomic data revealed that high levels of genomic differentiation are maintained across the majority of the genome when gene flow occurs. However despite this, we identified multiple, small regions of introgression, strongly correlated with recombination rate. Our results demonstrate that a high level of genome-wide divergence can establish in the face of persistent introgression and that gene flow can be localized to small genomic regions at the later stages of speciation with gene flow.\n\nAuthor summaryWhen species evolve, reproductive isolation leads to a build-up of differentiation in the genome where genes involved in the process occur. Much of our understanding of this comes from early stage speciation, with relatively few examples from more divergent species pairs that still exchange genes. To address this, we focused on Pacific Ocean and Japan Sea sticklebacks, which co-occur in the Japanese islands. We established that they are the oldest and most divergent known stickleback species pair, that they evolved in the face of gene flow and that this gene flow is still on going. We found introgression is confined to small, localised genomic regions where recombination rate is high. Our results show high divergence can be maintained between species, despite extensive gene flow.

evolutionary biology

Fast and Accurate Genomic Analyses using Genome Graphs

The human reference genome serves as the foundation for genomics by providing a scaffold for alignment of sequencing reads, but currently only reflects a single consensus haplotype, which impairs read alignment and downstream analysis accuracy. Reference genome structures incorporating known genetic variation have been shown to improve the accuracy of genomic analyses, but have so far remained computationally prohibitive for routine large-scale use. Here we present a graph genome implementation that enables read alignment across 2,800 diploid genomes encompassing 12.6 million SNPs and 4.0 million indels. Our Graph Genome Pipeline requires 6.5 hours to process a 30x coverage WGS sample on a system with 36 CPU cores compared with 11 hours required by the GATK Best Practices pipeline. Using complementary benchmarking experiments based on real and simulated data, we show that using a graph genome reference improves read mapping sensitivity and produces a 0.5% increase in variant calling recall, or about 20,000 additional variants being detected per sample, while variant calling specificity is unaffected. Structural variations (SVs) incorporated into a graph genome can be genotyped accurately under a unified framework. Finally, we show that iterative augmentation of graph genomes yields incremental gains in variant calling accuracy. Our implementation is a significant advance towards fulfilling the promise of graph genomes to radically enhance the scalability and accuracy of genomic analyses.

bioinformatics

Proteomics as a metrological tool to evaluate genome annotation accuracy following de novo genome assembly: a case study using the Atlantic bottlenose dolphin (Tursiops truncatus)

BackgroundThe last decade has witnessed dramatic improvements in whole-genome sequencing capabilities coupled to drastically decreased costs, leading to an inundation of high-quality de novo genomes. For this reason, continued development of genome quality metrics is imperative. The current study utilized the recently updated Atlantic bottlenose dolphin (Tursiops truncatus) genome and annotation to evaluate a proteomics-based metric of genome accuracy.\n\nResultsProteomic analysis of six tissues provided experimental confirmation of 10 402 proteins from 4 711 protein groups, almost 1/3 of the possible predicted proteins in the genome. There was an increased median molecular weight and number of identified peptides per protein using the current T. truncatus annotation versus the previous annotation. Identification of larger proteins with more identified peptides implied reduced database fragmentation and improved gene annotation accuracy. A metric is proposed, NP10, that attempts to capture this quality improvement. When using the new T. truncatus genome there was a 21 % improvement in NP10. This metric was further demonstrated by using a publicly available proteomic data set to compare human genome annotations from 2004, 2013 and 2016, which had a 33 % improvement in NP10.\n\nConclusionsThese results demonstrate that new whole-genome sequencing techniques can rapidly generate high quality de novo genome assemblies and emphasizes the speed of advancing bioanalytical measurements in a non-model organism. Moreover, proteomics may be a useful metrological tool to benchmark genome accuracy, though there is a need for reference proteomic datasets to facilitate this utility in new de novo and existing genomes.

bioinformatics

Fusobacterium genomics using MinION and Illumina sequencing enables genome completion and correction

Understanding the virulence mechanisms of human pathogens from the genus Fusobacterium has been hindered by a lack of properly assembled and annotated genomes. Here we report the first complete genomes for seven Fusobacterium strains, as well as resequencing of the reference strain F. nucleatum subsp. nucleatum ATCC 25586 (seven total species, eight total genomes). A highly efficient and cost-effective sequencing pipeline was achieved using sample multiplexing for short-read Illumina (150 bp) and long-read Oxford Nanopore MinION (>80 kbp) platforms, coupled with genome assembly using the open-source software Unicycler. When compared to currently available draft assemblies (previously 24-67 contigs), these genomes are highly accurate and consist of only one complete chromosome. We present the complete genome sequence of F. nucleatum 23726, a genetically tractable and biomedically important strain, and in addition, reveal that the previous F. nucleatum 25586 genome assembly contains a 452 kb genomic inversion that has been corrected using our sequencing and assembly pipeline. To enable the scientific community, we concurrently use these genomes to launch FusoPortal, a repository of interactive and downloadable genomic data, genome maps, gene annotations, and protein functional analysis and classification. In summary, this study provides detailed methods for accurately sequencing, assembling, and annotating Fusobacterium genomes, which will enhance efforts to properly identify virulence proteins that may contribute to a repertoire of diseases including periodontitis, pre-term birth, and colorectal cancer.

microbiology

A Thousand Fly Genomes: An Expanded Drosophila Genome Nexus

The Drosophila Genome Nexus is a population genomic resource that provides D. melanogaster genomes from multiple sources. To facilitate comparisons across data sets, genomes are aligned using a common reference alignment pipeline which involves two rounds of mapping. Regions of residual heterozygosity, identity-by-descent, and recent population admixture are annotated to enable data filtering based on the users needs. Here, we present a significant expansion of the Drosophila Genome Nexus, which brings the current data object to a total of 1,122 wild-derived genomes. New additions include 306 previously unpublished genomes from inbred lines representing six population samples in Egypt, Ethiopia, France, and South Africa, along with another 193 genomes added from recently-published data sets. We also provide an aligned D. simulans genome to facilitate divergence comparisons. This improved resource will broaden the range of population genomic questions that can addressed from multi-population allele frequencies and haplotypes in this model species. The larger set of genomes will also enhance the discovery of functionally relevant natural variation that exists within and between populations.

Genomics

Comparative Genomics and Genome Evolution in Birds-of-paradise

BackgroundThe diverse array of phenotypes and lekking behaviors in birds-of-paradise have long excited scientists and laymen alike. Remarkably, almost nothing is known about the genomics underlying this iconic radiation. Currently, there are 41 recognized species of birds-of-paradise, most of which live on the islands of New Guinea. In this study we sequenced genomes of representatives from all five major clades recognized within the birds-of-paradise family (Paradisaeidae). Our aim was to characterize genomic changes that may have been important for the evolution of the groups extensive phenotypic diversity.\n\nResultsWe sequenced three de novo genomes and re-sequenced two additional genomes representing all major clades within the birds-of-paradise. We found genes important for coloration, morphology and feather development to be under positive selection. GO enrichment of positively selected genes on the branch leading to the birds-of-paradise shows an enrichment for collagen, glycogen synthesis and regulation, eye development and other categories. In the core birds-of-paradise, we found GO categories for startle response (response to predators) and olfactory receptor activity to be enriched among the gene families expanding significantly faster compared to the other birds in our study. Furthermore, we found novel families of retrovirus-like retrotransposons active in all three de novo genomes since the early diversification of the birds-of-paradise group, which could have potentially played a role in the evolution of this fascinating group of birds.\n\nConclusionHere we provide a first glimpse into the genomic changes underlying the evolution of birds-of-paradise. Our aim was to use comparative genomics to study to what degree the genomic landscape of birds-of-paradise deviates from other closely related passerine birds. Given the extreme phenotypic diversity in this family, our prediction was that genomes should be able to reveal features important for the evolution of this amazing radiation. Overall, we found a strong signal for evolution on mechanisms important for coloration, morphology, sensory systems, as well as genome structure.

genomics

Resource: Scalable whole genome sequencing of 40,000 single cells identifies stochastic aneuploidies, genome replication states and clonal repertoires

Essential features of cancer tissue cellular heterogeneity such as negatively selected genome topologies, sub-clonal mutation patterns and genome replication states can only effectively be studied by sequencing single-cell genomes at scale and high fidelity. Using an amplification-free single-cell genome sequencing approach implemented on commodity hardware (DLP+) coupled with a cloud-based computational platform, we define a resource of 40,000 single-cell genomes characterized by their genome states, across a wide range of tissue types and conditions. We show that shallow sequencing across thousands of genomes permits reconstruction of clonal genomes to single nucleotide resolution through aggregation analysis of cells sharing higher order genome structure. From large-scale population analysis over thousands of cells, we identify rare cells exhibiting mitotic mis-segregation of whole chromosomes. We observe that tissue derived scWGS libraries exhibit lower rates of whole chromosome anueploidy than cell lines, and loss of p53 results in a shift in event type, but not overall prevalence in breast epithelium. Finally, we demonstrate that the replication states of genomes can be identified, allowing the number and proportion of replicating cells, as well as the chromosomal pattern of replication to be unambiguously identified in single-cell genome sequencing experiments. The combined annotated resource and approach provide a re-implementable large scale platform for studying lineages and tissue heterogeneity.

genomics

in silico Whole Genome Sequencer & Analyzer (iWGS): a computational pipeline to guide the design and analysis of de novo genome sequencing studies

The availability of genomes across the tree of life is highly biased toward vertebrates, pathogens, human disease models, and organisms with relatively small and simple genomes. Recent progress in genomics has enabled the de novo decoding of the genome of virtually any organism, greatly expanding its potential for understanding the biology and evolution of the full spectrum of biodiversity. The increasing diversity of sequencing technologies, assays, and de novo assembly algorithms have augmented the complexity of de novo genome sequencing projects in non-model organisms. To reduce the costs and challenges in de novo genome sequencing projects and streamline their experimental design and analysis, we developed iWGS (in silico Whole Genome Sequencer and Analyzer), an automated pipeline for guiding the choice of appropriate sequencing strategy and assembly protocols. iWGS seamlessly integrates the four key steps of a de novo genome sequencing project: data generation (through simulation), data quality control, de novo assembly, and assembly evaluation and validation. The last three steps can also be applied to the analysis of real data. iWGS is designed to enable the user to have great flexibility in testing the range of experimental designs available for genome sequencing projects, and supports all major sequencing technologies and popular assembly tools. Three case studies illustrate how iWGS can guide the design of de novo genome sequencing projects and evaluate the performance of a wide variety of user-specified sequencing strategies and assembly protocols on genomes of differing architectures. iWGS, along with a detailed documentation, is freely available at https://github.com/zhouxiaofan1983/iWGS.

Bioinformatics

Timing and scope of genomic expansion within Annelida: evidence from homeoboxes in the genome of the earthworm Eisenia fetida

Annelida represents a large and morphologically diverse group of bilaterian organisms. The recently published polychaete and leech genome sequences revealed an equally dynamic range of diversity at the genomic level. The availability of more annelid genomes will allow for the identification of evolutionary genomic events that helped shape the annelid lineage and better understand the diversity within the group. We sequenced and assembled the genome of the common earthworm, Eisenia fetida. As a first pass at understanding the diversity within the group, we classified 440 earthworm homeoboxes and compared them to those of the leech Helobdella robusta and the polychaete Capitella teleta. We inferred many gene expansions occurring in the lineage connecting the most recent common ancestor (MRCA) of Capitella and Eisenia to the Eisenia/Helobdella MRCA. Likewise, the lineage leading from the Eisenia/Helobdella MRCA to the leech Helobdella robusta has experienced substantial gains and losses. However, the lineage leading from Eisenia/Helobdella MRCA to E. fetida is characterized by extraordinary levels of homeobox gain. The evolutionary dynamics observed in the homeoboxes of these lineages are very likely to be generalizable to all genes. These genome expansions and losses have likely contributed to the remarkable biology exhibited in this group. These results provide a new perspective from which to understand the diversity within these lineages, show the utility of sub-draft genome assemblies for understanding genomic evolution, and provide a critical resource from which the biology of these animals can be studied. The genome data can be accessed through the Eisenia fetida Genome Portal: http://ryanlab.whitney.ufl.edu/genomes/Efet/

Genomics

Hi-C deconvolution of a human gut microbiome yields high-quality draft genomes and reveals plasmid-genome interactions.

The assembly of high-quality genomes from mixed microbial samples is a long-standing challenge in genomics and metagenomics. Here, we describe the application of ProxiMeta, a Hi-C-based metagenomic deconvolution method, to deconvolve a human fecal metagenome. This method uses the intra-cellular proximity signal captured by Hi-C as a direct indicator of which sequences originated in the same cell, enabling culture-free de novo deconvolution of mixed genomes without any reliance on a priori information. We show that ProxiMeta deconvolution provides results of markedly high accuracy and sensitivity, yielding 50 near-complete microbial genomes (many of which are novel) from a single fecal sample, out of 252 total genome clusters. ProxiMeta outperforms traditional contig binning at high-quality genome reconstruction. ProxiMeta shows particularly good performance in constructing high-quality genomes for diverse but poorly-characterized members of the human gut. We further use ProxiMeta to reconstruct genome plasmid content and sharing of plasmids among genomes--tasks that traditional binning methods usually fail to accomplish. Our findings suggest that Hi-C-based deconvolution can be useful to a variety of applications in genomics and metagenomics.

genomics

Pushing the limits of de novo genome assembly for complex prokaryotic genomes harboring very long, near identical repeats

Generating a complete, de novo genome assembly for prokaryotes is often considered a solved problem. However, we here show that Pseudomonas koreensis P19E3 harbors multiple, near identical repeat pairs up to 70 kilobase pairs in length. Beyond long repeats, the P19E3 assembly was further complicated by a shufflon region. Its complex genome could not be de novo assembled with long reads produced by Pacific Biosciences technology, but required very long reads from the Oxford Nanopore Technology. Another important factor for a full genomic resolution was the choice of assembly algorithm.\n\nImportantly, a repeat analysis indicated that very complex bacterial genomes represent a general phenomenon beyond Pseudomonas. Roughly 10% of 9331 complete bacterial and a handful of 293 complete archaeal genomes represented this dark matter for de novo genome assembly of prokaryotes. Several of these dark matter genome assemblies contained repeats far beyond the resolution of the sequencing technology employed and likely contain errors, other genomes were closed employing labor-intense steps like cosmid libraries, primer walking or optical mapping. Using very long sequencing reads in combination with assemblers capable of resolving long, near identical repeats will bring most prokaryotic genomes within reach of fast and complete de novo genome assembly.

genomics