Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Full-genome evolutionary histories of selfing, splitting and selection in Caenorhabditis

The nematode Caenorhabditis briggsae is a model for comparative developmental evolution with C. elegans. Worldwide collections of C. briggsae have implicated an intriguing history of divergence among genetic groups separated by latitude, or by restricted geography, that is being exploited to dissect the genetic basis to adaptive evolution and reproductive incompatibility. And yet, the genomic scope and timing of population divergence is unclear. We performed high-coverage whole-genome sequencing of 37 wild isolates of the nematode C. briggsae and applied a pairwise sequentially Markovian coalescent (PSMC) model to 703 combinations of genomic haplotypes to draw inferences about population history, the genomic scope of natural selection, and to compare with 40 wild isolates of C. elegans. We estimate that a diaspora of at least 6 distinct C. briggsae lineages separated from one another approximately 200 thousand generations ago, including the Temperate and Tropical phylogeographic groups that dominate most samples from around the world. Moreover, an ancient population split in its history 2 million generations ago, coupled with only rare gene flow among lineage groups, validates this system as a model for incipient speciation. Low versus high recombination regions of the genome give distinct signatures of population size change through time, indicative of widespread effects of selection on highly linked portions of the genome owing to extreme inbreeding by self-fertilization. Analysis of functional mutations indicates that genomic context, owing to selection that acts on long linkage blocks, is a more important driver of population variation than are the functional attributes of the individually encoded genes.

Genomics

Efficient Privacy-Preserving String Search and an Application in Genomics

Motivation: Personal genomes carry inherent privacy risks and protecting privacy poses major social and technological challenges. We consider the case where a user searches for genetic information (e.g., an allele) on a server that stores a large genomic database and aims to receive allele-associated information. The user would like to keep the query and result private and the server the database.\n\nApproach: We propose a novel approach that combines efficient string data structures such as the Burrows-Wheeler transform with cryptographic techniques based on additive homomorphic encryption. We assume that the sequence data is searchable in efficient iterative query operations over a large indexed dictionary, for instance, from large genome collections and employing the (positional) Burrows-Wheeler transform. We use a technique called oblivious transfer that is based on additive homomorphic encryption to conceal the sequence query and the genomic region of interest in positional queries.\n\nResults: We designed and implemented an efficient algorithm for searching sequences of SNPs in large genome databases. During search, the user can only identify the longest match while the server does not learn which sequence of SNPs the user queried. In an experiment based on 2,184 aligned haploid genomes from the 1,000 Genomes Project, our algorithm was able to perform typical queries within {approx}4.6 seconds and {approx}10.8 seconds for client and server side, respectively, on laptop computers. The presented algorithm is at least one order of magnitude faster than an exhaustive baseline algorithm.\n\nAvailability: https://github.com/iskana/PBWT-sec and https://github.com/ratschlab/PBWT-sec.

Genomics

Sequencing of 15,622 gene-bearing BACs reveals new features of the barley genome

Barley (Hordeum vulgare L.) possesses a large and highly repetitive genome of 5.1 Gb that has hindered the development of a complete sequence. In 2012, the International Barley Sequencing Consortium released a resource integrating whole-genome shotgun sequences with a physical and genetic framework. However, since only 6,278 BACs in the physical map were sequenced, detailed fine structure was limited. To gain access to the gene-containing portion of the barley genome at high resolution, we identified and sequenced 15,622 BACs representing the minimal tiling path of 72,052 physical mapped gene-bearing BACs. This generated about 1.7 Gb of genomic sequence containing 17,386 annotated barley genes. Exploration of the sequenced BACs revealed that although distal ends of chromosomes contain most of the gene-enriched BACs and are characterized by high rates of recombination, there are also gene-dense regions with suppressed recombination. Knowledge of these deviant regions is relevant to trait introgression, genome-wide association studies, genomic selection model development and map-based cloning strategies. Sequences and their gene and SNP annotations can be accessed and exported via http://harvest-web.org/hweb/utilmenu.wc or through the software HarvEST:Barley (download from harvest.ucr.edu). In the latter, we have implemented a synteny viewer between barley and Aegilops tauschii to aid in comparative genome analysis.

Genomics

Contiguous and accurate de novo assembly of metazoan genomes with modest long read coverage

Genome assemblies that are accurate, complete, and contiguous are essential for identifying important structural and functional elements of genomes and for identifying genetic variation. Nevertheless, most recent genome assemblies remain incomplete and fragmented. While long molecule sequencing promises to deliver more complete genome assemblies with fewer gaps, concerns about error rates, low yields, stringent DNA requirements, and uncertainty about best practices may discourage many investigators from adopting this technology. Here, in conjunction with the platinum standard Drosophila melanogaster reference genome, we analyze recently published long molecule sequencing data to identify what governs completeness and contiguity of genome assemblies. We also present a hybrid meta-assembly approach that achieves remarkable assembly contiguity for both Drosophila and human assemblies with only modest long molecule sequencing coverage. Our results motivate a set of preliminary best practices for obtaining accurate and contiguous assemblies, a \"missing manual\" that guides key decisions in building high quality de novo genome assemblies, from DNA isolation to polishing the assembly.

Genomics

MaGuS: a tool for map-guided scaffolding and quality assessment of genome assemblies

BackgroundScaffolding is a crucial step in the genome assembly process. Current methods based on large fragment paired-end reads or long reads allow an increase in continuity but often lack consistency in repetitive regions, resulting in fragmented assemblies. Here, we describe a novel tool to link assemblies to a genome map to aid complex genome reconstruction by detecting assembly errors and allowing scaffold ordering and anchoring.\n\nResultsWe present MaGuS (map-guided scaffolding), a modular tool that uses a draft genome assembly, a genome map, and high-throughput paired-end sequencing data to estimate the quality and to enhance the continuity of an assembly. We generated several assemblies of the Arabidopsis genome using different scaffolding programs and applied MaGuS to select the best assembly using quality metrics. Then, we used MaGuS to perform map-guided scaffolding to increase continuity by creating new scaffold links in low-covered and highly repetitive regions where other commonly used scaffolding methods lack consistency.\n\nConclusionsMaGuS is a powerful reference-free evaluator of assembly quality and a map-guided scaffolder that is freely available at https://github.com/institut-de-genomique/MaGuS. Its use can be extended to other high-throughput sequencing data (e.g., long-read data) and also to other map data (e.g., genetic maps) to improve the quality and the continuity of large and complex genome assemblies.

Genomics

Improvement of the threespine stickleback (Gasterosteus aculeatus) genome using a Hi-C-based Proximity-Guided Assembly method

Scaffolding genomes into complete chromosome assemblies remains challenging even with the rapidly increasing sequence coverage generated by current next-generation sequence technologies. Even with scaffolding information, many genome assemblies remain incomplete. The genome of the threespine stickleback (Gasterosteus aculeatus), a fish model system in evolutionary genetics and genomics, is not completely assembled despite scaffolding with high-density linkage maps. Here, we first test the ability of a Hi-C based proximity guided assembly to perform a de novo genome assembly from relatively short contigs. Using Hi-C based proximity guided assembly, we generated complete chromosome assemblies from 50 kb contigs. We found that 98.99% of contigs were correctly assigned to linkage groups, with ordering nearly identical to the previous genome assembly. Using available BAC end sequences, we provide evidence that some of the few discrepancies between the Hi-C assembly and the existing assembly are due to structural variation between the populations used for the two assemblies or errors in the existing assembly. This Hi-C assembly also allowed us to improve the existing assembly, assigning over 60% (13.35 Mb) of the previously unassigned ([~]21.7 Mb) contigs to linkage groups. Together, our results highlight the potential of the Hi-C based proximity guided assembly method to be used in combination with short read data to perform relatively inexpensive de novo genome assemblies. This approach will be particularly useful in organisms in which it is difficult to perform linkage mapping or to obtain high molecular weight DNA required for other scaffolding methods.

Genomics

Using RNA-seq for genomic scaffold placement, correcting assemblies, and genetic map creation in a common Brassica rapa mapping population.

Brassica rapa is a model species for agronomic, ecological, evolutionary and translational studies. Here we describe high-density SNP discovery and genetic map construction for a Brassica rapa recombinant inbred line (RIL) population derived from field collected RNA-seq data. This high-density genotype data enables the detection and correction of putative genome mis-assemblies and accurate assignment of scaffold sequences to their likely genomic locations. These assembly improvements represent 7.1-8.0% of the annotated Brassica rapa genome. We demonstrate how using this new resource leads to a significant improvement for QTL analysis over the current low-density genetic map. Improvements are achieved by the increased mapping resolution and by having known genomic coordinates to anchor the markers for candidate gene discovery. These new molecular resources and improvements in the genome annotation will benefit the Brassicaceae genomics community and may help guide other communities in finetuning genome annotations.

Genomics

Nanopore DNA Sequencing and Genome Assembly on the International Space Station

The emergence of nanopore-based sequencers greatly expands the reach of sequencing into low-resource field environments, enabling in situ molecular analysis. In this work, we evaluated the performance of the MinION DNA sequencer (Oxford Nanopore Technologies) in-flight on the International Space Station (ISS), and benchmarked its performance off-Earth against the MinION, Illumina MiSeq, and PacBio RS II sequencing platforms in terrestrial laboratories. Samples contained mixtures of genomic DNA extracted from lambda bacteriophage, Escherichia coli (strain K12) and Mus musculus (BALB/c). The in-flight sequencing experiments generated more than 80,000 total reads with mean 2D accuracies of 85 - 90%, mean 1D accuracies of 75 - 80%, and median read lengths of approximately 6,000 bases. We were able to construct directed assemblies of the ~4.7 Mb E. coli genome, ~48.5 kb lambda genome, and a representative M. musculus sequence (the ~16.3 kb mitochondrial genome), at 100%, 100%, and 96.7% pairwise identity, respectively, and de novo assemblies of the lambda and E. coli genomes generated solely from nanopore reads yielded 100% and 99.8% genome coverage, respectively, at 100% and 98.5% pairwise identity. Across all surveyed metrics (base quality, throughput, stays/base, skips/base), no observable decrease in MinION performance was observed while sequencing DNA in space. Simulated runs of in-flight nanopore data using an automated bioinformatic pipeline and cloud or laptop based genomic assembly demonstrated the feasibility of real-time sequencing analysis and direct microbial identification in space. Applications of sequencing for space exploration include infectious disease diagnosis, environmental monitoring, evaluating biological responses to spaceflight, and even potentially the detection of extraterrestrial life on other planetary bodies.

Genomics

Darwin: A Hardware-acceleration Framework for Genomic Sequence Alignment

Genomics is set to transform medicine and our understanding of life in fundamental ways. But the growth in genomics data has been overwhelming - far outpacing Moores Law. The advent of third generation sequencing technologies is providing new insights into genomic contribution to diseases with complex mutation events, but have prohibitively high computational costs. Over 1,300 CPU hours are required to align reads from a 54x coverage of the human genome to a reference (estimated using [1]), and over 15,600 CPU hours to assemble the reads de novo [2]. This paper proposes \"Darwin\" - a hardware-accelerated framework for genomic sequence alignment that, without sacrificing sensitivity, provides 125x and 15.6x speedup over the state-of-the-art software counterparts for reference-guided and de novo assembly of third generation sequencing reads, respectively. For pairwise alignment of sequences, Darwin is over 39,000x more energy-efficient than software. Darwin uses (i) a novel filtration strategy, called D-SOFT, to reduce the search space for sequence alignment at high speed, and (ii) a hardware-accelerated version of GACT, a novel algorithm to generate near-optimal alignments of arbitrarily long genomic sequences using constant memory for trace-back. Darwin is adaptable, with tunable speed and sensitivity to match emerging sequencing technologies and to meet the requirements of genomic applications beyond read assembly.

genomics

Rapid Automated Large Structural Variation Detection in a Diploid Genome by NanoChannel Based Next-Generation Mapping

The human genome is diploid with one haploid genome inherited from the maternal and one from the paternal lineage. Within each haploid genome, large structural variants such as deletions, duplications, inversions, and translocations are extensively present and many are known to affect biological functions and cause disease. The ultimate goal is to resolve these large complex structural variants (SVs) and place them in the correct haploid genome with correct location, orientation, and copy number. Current methods such as karyotyping, chromosomal microarray (CMA), PCR-based tests, and next-generation sequencing fail to reach this goal either due to limited resolution, low throughput, or short read length.\n\nBionano Genomics next-generation mapping (NGM) offers a high-throughput, genome-wide method able to detect SVs of one kilobase pairs (kbp) and up. By imaging extremely long genomic molecules of up to megabases in size, the structure and copy number of complex regions of the g ...

genomics

Mapping And Phasing Of Structural Variation In Patient Genomes Using Nanopore Sequencing

Structural genomic variants form a common type of genetic alteration underlying human genetic disease and phenotypic variation. Despite major improvements in genome sequencing technology and data analysis, the detection of structural variants still poses challenges, particularly when variants are of high complexity. Emerging long-read single-molecule sequencing technologies provide new opportunities for detection of structural variants. Here, we demonstrate sequencing of the genomes of two patients with congenital abnormalities using the ONT MinION at 11x and 16x mean coverage, respectively. We developed a bioinformatic pipeline - NanoSV - to efficiently map genomic structural variants (SVs) from the long-read data. We demonstrate that the nanopore data are superior to corresponding short-read data with regard to detection of de novo rearrangements originating from complex chromothripsis events in the patients. Additionally, genome-wide surveillance of SVs, revealed 3,253 (33%) novel variants that were missed in short-read data of the same sample, the majority of which are duplications < 200bp in size. Long sequencing reads enabled efficient phasing of genetic variations, allowing the construction of genome-wide maps of phased SVs and SNVs. We employed read-based phasing to show that all de novo chromothripsis breakpoints occurred on paternal chromosomes and we resolved the long-range structure of the chromothripsis. This work demonstrates the value of long-read sequencing for screening whole genomes of patients for complex structural variants.

genomics

Ultra-Accurate Genome Sequencing And Haplotyping Of Single Human Cells

Accurate detection of variants and long-range haplotypes in genomes of single human cells remains very challenging. Common approaches require extensive in vitro amplification of genomes of individual cells using DNA polymerases and high-throughput short-read DNA sequencing. These approaches have two notable drawbacks. First, polymerase replication errors could generate tens of thousands of false positive calls per genome. Second, relatively short sequence reads contain little to no haplotype information. Here we report a method, which is dubbed SISSOR (Single-Stranded Sequencing using micrOfluidic Reactors), for accurate single-cell genome sequencing and haplotyping. A microfluidic processor is used to separate the Watson and Crick strands of the double-stranded chromosomal DNA in a single cell and to randomly partition megabase-size DNA strands into multiple nanoliter compartments for amplification and construction of barcoded libraries for sequencing. The separation and partitioning of large single-stranded DNA fragments of the homologous chromosome pairs allows for the independent sequencing of each of the complementary and homologous strands. This enables the assembly of long haplotypes and reduction of sequence errors by using the redundant sequence information and haplotype-based error removal. We demonstrated the ability to sequence single-cell genomes with error rates as low as 10-8 and average 500kb long DNA fragments that can be assembled into haplotype contigs with N50 greater than 7Mb. The performance could be further improved with more uniform amplification and more accurate sequence alignment. The ability to obtain accurate genome sequences and haplotype information from single cells will enable applications of genome sequencing for diverse clinical needs.

genomics

The Drosophila Y chromosome affects heterochromatin integrity genome-wide

The Drosophila Y-chromosome is gene poor and mainly consists of silenced, repetitive DNA. Nonetheless, the Y influences expression of hundreds of genes genome-wide, possibly by sequestering key components of the heterochromatin machinery away from other positions in the genome. To test the influence of the Y chromosome on the genome-wide chromatin landscape, we assayed the genomic distribution of histone modifications associated with gene activation (H3K4me3), or heterochromatin (H3K9me2 and H3K9me3) in fruit flies with varying sex chromosome complements (X0, XY and XYY males; XX and XXY females). Consistent with the general deficiency of active chromatin modifications on the Y, we find that Y gene dose has little influence on the genomic distribution of H3K4me3. In contrast, both the presence and the number of Y-chromosomes strongly influence genome-wide enrichment patterns of repressive chromatin modifications. Highly repetitive regions such as the pericentromeres, the dot, and the Y chromosome (if present) are enriched for heterochromatic modifications in wildtype males and females, and even more strongly in X0 flies. In contrast, the additional Y chromosome in XYY males and XXY females diminishes the heterochromatic signal in these normally silenced, repeat-rich regions, which is accompanied by an increase in expression of Y-linked repeats. We find hundreds of genes that are expressed differentially between individuals with aberrant sex chromosome karyotypes, many of which also show sex-biased expression in wildtype Drosophila. Thus, Y-chromosomes influence heterochromatin integrity genome-wide, and differences in the chromatin landscape of males and females may also contribute to sex-biased gene expression and sexual dimorphisms.

genomics

Whole genome optical mapping reveals multiple fusion events chained by large novel sequences in cancer

Genomic rearrangements are common in cancer, with demonstrated links to disease progression and treatment response. These rearrangements can be complex, resulting in fusions of multiple chromosomal fragments and generation of derivative chromosomes. While methods exist for detecting individual fusions, they are generally unable to reconstruct complex chained events. To overcome these limitations, we adopted a new optical mapping approach, allowing for megabase length DNA to be captured, and in turn rearranged genomes to be visualized without loss of integrity. Whole genome mapping (Bionano Genomics) of a well-studied highly rearranged liposarcoma cell line, resulted in 3,338 assembled haploid genome maps, including 101 fusion maps. These fusion maps represent 175 Mb of highly rearranged genomic regions, illuminating the complex architecture of chained fusions, including content, order, orientation, and size. Spanning the junction of 151 chromosomal translocations, we found a total of 32 Mb of novel interspersed sequences that were not detected from short-read sequencing. We demonstrate that optical mapping provides a powerful new approach for capturing a higher level of complex genomic architecture, creating a scaffold for renewed interpretation of sequencing data of particular relevance to human cancer.

genomics

In situ genome editing method suitable for routine generation of germline modified animal models

Animal genome engineering experimental procedures involve three major steps: isolation of zygotes from pregnant females; microinjection of zygotes, and; transfer of injected zygotes into recipient females, that have been practiced for over three decades. The laboratory set ups intending to performing these procedures require to have sophisticated equipment as well as highly skilled technical personnel. Because of these reasons, animal transgenesis experiments are typically performed at centralized core facilities in most research organizations. We recently showed that all three steps, of animal transgensis, can be bypassed using a method termed GONAD (Genome-editing via Oviductal Nucleic Acids Delivery), by directly electroporating genome editing components into zygotes in situ. Although our first report demonstrated the genome-editing capability, its efficiency was lower than the standard methods using microinjection. Here we investigated critical parameters of GONAD to make it suitable for creating animal models of large genomic deletions, single nucleotide corrections and long sequence insertions. The efficiency of genome editing in the improved GONAD (i-GONAD) method reached to the levels comparable to traditional microinjection methods. The streamlined parameters, and the simplified experimental steps, in the i-GONAD method makes it suitable for routine genome editing applications performed both at centralized facilities as well as at the laboratories that lack highly skilled personnel and the sophisticated equipment.

genomics

SuperDCA for genome-wide epistasis analysis

The potential for genome-wide modeling of epistasis has recently surfaced given the possibility of sequencing densely sampled populations and the emerging families of statistical interaction models. Direct coupling analysis (DCA) has earlier been shown to yield valuable predictions for single protein structures, and has recently been extended to genome-wide analysis of bacteria, identifying novel interactions in the co-evolution between resistance, virulence and core genome elements. However, earlier computational DCA methods have not been scalable to enable model fitting simultaneously to 104-105 polymorphisms, representing the amount of core genomic variation observed in analyses of many bacterial species. Here we introduce a novel inference method (SuperDCA) which employs a new scoring principle, efficient parallelization, optimization and filtering on phylogenetic information to achieve scalability for up to 105 polymorphisms. Using two large population samples of Streptococcus pneumoniae, we demonstrate the ability of SuperDCA to make additional significant biological findings about this major human pathogen. We also show that our method can uncover signals of selection that are not detectable by genome-wide association analysis, even though our analysis does not require phenotypic measurements. SuperDCA thus holds considerable potential in building understanding about numerous organisms at a systems biological level.\n\nAuthor SummaryRecent work has demonstrated the emerging potential in statistical genome-wide modeling to uncover co-selection and epistatic interactions between polymorphisms in bacterial chromosomes from densely sampled population data. Here we develop the Potts model based approach further into a fully mature computational method which can be applied to most existing bacterial population genomic data sets in a straightforward manner. Our advances are relying on more efficient parameter scoring, highly optimized and parallelized open source C++ code, which does not rely on the computation-intensive polymorphism subsampling approximations used earlier. By analyzing the two largest available population samples of Streptococcus pneumoniae (the pneumococcus), we highlight several biological discoveries related to the survival of the pneumococcus and co-evolution of penicillin-binding loci, which were not uncovered by the earlier analyses. Our method holds considerable potential for building understanding about numerous organisms at a systems biological level.

genomics

Interoperable genome annotation with GBOL, an extendable infrastructure for functional data mining

BackgroundA standard structured format is used by the public sequence databases to present genome annotations. A prerequisite for a direct functional comparison is consistent annotation of the genetic elements with evidence statements. However, the current format provides limited support for data mining, hampering comparative analyses at large scale.\n\nResultsThe provenance of a genome annotation describes the contextual details and derivation history of the process that resulted in the annotation. To enable interoperability of genome annotations, we have developed the Genome Biology Ontology Language (GBOL) and associated infrastructure (GBOL stack). GBOL is provenance aware and thus provides a consistent representation of functional genome annotations linked to the provenance. GBOL is modular in design, extendible and linked to existing ontologies. The GBOL stack of supporting tools enforces consistency within and between the GBOL definitions in the ontology (OWL) and the Shape Expressions (ShEx) language describing the graph structure. Modules have been developed to serialize the linked data (RDF) and to generate a plain text format files.\n\nConclusionThe main rationale for applying formalized information models is to improve the exchange of information. GBOL uses and extends current ontologies to provide a formal representation of genomic entities, along with their properties and relations. The deliberate integration of data provenance in the ontology enables review of automatically obtained genome annotations at a large scale. The GBOL stack facilitates consistent usage of the ontology.

genomics

Comparison of seven single cell Whole Genome Amplification commercial kits using targeted sequencing

Advances in biochemical technologies have led to a boost in the field of single cell genomics. Observation of the genome at a single cell resolution is currently achieved by pre-amplification using whole genome amplification (WGA) techniques that differ by their biochemical aspects and as a result by biased amplification of the original molecule. Several comparisons between commercially available single cell dedicated WGA kits (scWGA) were performed, however, these comparisons are costly, were only performed on selected scWGA kit and more notably, are limited by the number of analyzed cells, making them limited for reproducibility analysis. We benchmarked an economical assay to compare all commercially available scWGA kits that is based on targeted sequencing of thousands of genomic regions, including highly mutable genomic regions (microsatellites), from a large cohort of human single cells (125 cells in total). Using this approach, we could analyze the genome coverage, the reproducibility of genome coverage and the error rate of each kit. Our experimental design provides an affordable and reliable comparative assay that simulates a real single cell experiment. Results demonstrate the needfor a dedicated kit selection depending on the desired single cell assay.

genomics