Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

A Novel Method for the Capture-based Purification of Whole Viral Native RNA Genomes

Current technologies for targeted characterization and manipulation of viral RNA either involve amplification or ultracentrifugation with isopycnic gradients of viral particles to decrease host RNA background. The former strategy is non-compatible for characterizing properties innate to RNA strands such as secondary structure, RNA-RNA interactions, and also for nanopore direct RNA sequencing involving the sequencing of native RNA strands. The latter strategy, ultracentrifugation, causes loss in genomic information due to its inability to retrieve unassembled viral RNA. We developed a novel nucleic acid manipulation technique involving the capture of whole viral native RNA genomes for downstream RNA assays using hybridization baits in solution to circumvent these problems. This technique involves hybridization of biotinylated baits at 500 nucleotides (nt) intervals, stringent washes and release of free native RNA strands using DNase I treatment, with a turnaround time of about 6 h 15 min. Proof of concept was primarily done using RT-qPCR with dengue virus infected Huh-7 cells. We report that this protocol was able to purify viral RNA (561-791 fold). We also describe a successful application of our capture-based purification method to direct RNA sequencing, with a 77.47% of reads mapping to the target viral genome. We observed a reduction in human host RNA background by 1580 fold, a 99.91% recovery of viral genome with at least 15x coverage, and a mean coverage across the genome of 120x. This report is, to the best of our knowledge, the first description of a capture-based purification method for whole viral RNA genomes. The fundamental advantages of using our capture-based purification method makes it a superior alternative to conventional viral purification methods and would potentially pave a new path for the direct characterization and sequencing of native RNA molecules.

genomics

Genome-wide selection footprints and deleterious variations in young Asian allotetraploid rapeseed

Brassica napus (AACC, 2n=38), is an important oilseed crop grown worldwide. However, little is known about the population evolution of this species, the genomic difference between its major genetic clusters, such as European and Asian rapeseed, and impacts of historical large-sale introgression events in this young tetraploid. In this study, we reported the de novo assembly of the genome sequences of an Asian rapeseed (B. napus), Ningyou 7 and its four progenitors and carried out de novo assembly-based comparison, pedigree and population analysis with other available genomic data from diverse European and Asian cultivars. Our results showed that Asian rapeseed originally derived from European rapeseed, but it had subsequently significantly diverged, with rapid genome differentiation after intensive local breeding selection. The first historical introgression of B. rapa dramatically broadened the allelic pool of Asian B. napus, but decreased their deleterious variations. The secondary historical introgression of European rapeseed (canola-quality) has reshaped Asian rapeseed into two groups, accompanied by an increase in genetic load. This study demonstrates distinctive genomic footprints by recent intra- and inter-species introgression events for local adaptation, and provide novel insights for understanding the rapid genome evolution of a young allopolyploid crop.

genomics

One Health genomic surveillance of Escherichia coli demonstrates distinct lineages and mobile genetic elements in isolates from humans versus livestock

Livestock have been proposed as a reservoir for drug-resistant Escherichia coli that infect humans. We isolated and sequenced 431 E. coli (including 155 ESBL-producing isolates) from cross-sectional surveys of livestock farms and retail meat in the East of England. These were compared with the genomes of 1517 E. coli associated with bloodstream infection in the United Kingdom. Phylogenetic core genome comparisons demonstrated that livestock and patient isolates were genetically distinct, indicating that E. coli causing serious human infection do not directly originate from livestock. By contrast, we observed highly related isolates from the same animal species on different farms. Analysis of accessory (variable) genomes identified a virulence cassette associated previously with cystitis and neonatal meningitis that was only present in isolates from humans. Screening all 1948 isolates for accessory genes encoding antibiotic resistance revealed 41 different genes present in variable proportions of humans and livestock isolates. We identified a low prevalence of shared antimicrobial resistance genes between livestock and humans based on analysis of mobile genetic elements and long-read sequencing. We conclude that in this setting, there was limited evidence to support the suggestion that antimicrobial resistant pathogens that cause serious infection in humans originate from livestock.\n\nImportanceThe increasing prevalence of E. coli bloodstream infections is a serious public health problem. We used genomic epidemiology in a One Health study conducted in the East of England to examine putative sources of E. coli associated with serious human disease. E. coli from 1517 patients with bloodstream infection were compared with 431 isolates from livestock farms and meat. Livestock-associated and bloodstream isolates were genetically distinct populations based on core genome and accessory genome analyses. Identical antimicrobial resistance genes were found in livestock and human isolates, but there was little overlap in the mobile elements carrying these genes. In addition, a virulence cassette found in humans isolates was not identified in any livestock-associated isolate. Our findings do not support the idea that E. coli causing invasive disease or their resistance genes are commonly acquired from livestock.

genomics

Genome sequence of the corn leaf aphid (Rhopalosiphum maidis Fitch)

BackgroundThe corn leaf aphid (Rhopalosiphum maidis Fitch) is the most economically damaging aphid pest on maize (Zea mays), one of the worlds most important grain crops. In addition to causing direct damage due to the removal of photoassimilates, R. maidis transmits several destructive maize viruses, including Maize yellow dwarf virus, Barley yellow dwarf virus, Sugarcane mosaic virus, and Cucumber mosaic virus.\n\nFindingsA 326-Mb genome assembly of BTI-1, a parthenogenetically reproducing R. maidis clone, was generated with a combination of PacBio (208-fold coverage) and Illumina sequencing (80-fold coverage), which contains a total of 689 contigs with an N50 size of 9.0 Mb. The contigs were further clustered into four scaffolds using the Phase Genomics Hi-C interaction maps, consistent with the commonly observed 2n = 8 karyotype of R. maidis. Most of the assembled contigs (473 spanning 321 Mb) were successfully orientated in the four scaffolds. The R. maidis genome assembly captured the full length of 95.8% of the core eukaryotic genes, suggesting that it is highly complete. Repetitive sequences accounted for 21.2% of the assembly, and a total of 17,647 protein-coding genes were predicted in the R. maidis genome with integrated evidence from ab initio and homology-based gene predictions and transcriptome sequences generated with both PacBio and Illumina. An analysis of likely horizontally transferred genes identified two from bacteria, seven from fungi, two from protozoa, and nine from algae.\n\nConclusionsA high-quality R. maidis genome was assembled at the chromosome level. This genome sequence will enable further research related to ecological interactions, virus transmission, pesticide resistance, and other aspects of R. maidis biology. It also serves as a valuable resource for comparative investigation of other aphid species.

genomics

The genome of the plague-resistant great gerbil reveals species-specific duplication of an MHCII gene

The great gerbil (Rhombomys opimus) is a social rodent living in permanent, complex burrow systems distributed throughout Central Asia, where it serves as the main host of several important vector-borne infectious diseases and is defined as a key reservoir species for plague (Yersinia pestis). Studies from the wild have shown that the great gerbil is largely resistant to plague but the genetic basis for resistance is yet to be determined. Here, we present a highly contiguous annotated genome assembly of great gerbil, covering over 96 % of the estimated 2.47 Gb genome. Comparative genomic analyses focusing on the immune gene repertoire, reveal shared gene losses within TLR gene families (i.e. TLR8, TLR10 and all members of TLR11-subfamily) for the Gerbillinae lineage, accompanied with signs of diversifying selection of TLR7 and TLR9. Most notably, we find a great gerbil-specific duplication of the MHCII DRB locus. In silico analyses suggest that the duplicated gene provides high peptide binding affinity for Yersiniae epitopes. The great gerbil genome provides new insights into the genomic landscape that confers immunological resistance towards plague. The high affinity for Yersinia epitopes could be key in our understanding of the high resistance in great gerbils, putatively conferring a faster initiation of the adaptive immune response leading to survival of the infection. Our study demonstrates the power of studying zoonosis in natural hosts through the generation of a genome resource for further comparative and experimental work on plague survival and evolution of host-pathogen interactions.

genomics

The genetics of the mood disorder spectrum: genome-wide association analyses of over 185,000 cases and 439,000 controls

BackgroundMood disorders (including major depressive disorder and bipolar disorder) affect 10-20% of the population. They range from brief, mild episodes to severe, incapacitating conditions that markedly impact lives. Despite their diagnostic distinction, multiple approaches have shown considerable sharing of risk factors across the mood disorders.\n\nMethodsTo clarify their shared molecular genetic basis, and to highlight disorder-specific associations, we meta-analysed data from the latest Psychiatric Genomics Consortium (PGC) genome-wide association studies of major depression (including data from 23andMe) and bipolar disorder, and an additional major depressive disorder cohort from UK Biobank (total: 185,285 cases, 439,741 controls; non-overlapping N = 609,424).\n\nResultsSeventy-three loci reached genome-wide significance in the meta-analysis, including 15 that are novel for mood disorders. More genome-wide significant loci from the PGC analysis of major depression than bipolar disorder reached genome-wide significance. Genetic correlations revealed that type 2 bipolar disorder correlates strongly with recurrent and single episode major depressive disorder. Systems biology analyses highlight both similarities and differences between the mood disorders, particularly in the mouse brain cell types implicated by the expression patterns of associated genes. The mood disorders also differ in their genetic correlation with educational attainment - positive in bipolar disorder but negative in major depressive disorder.\n\nConclusionsThe mood disorders share several genetic associations, and can be combined effectively to increase variant discovery. However, we demonstrate several differences between these disorders. Analysing subtypes of major depressive disorder and bipolar disorder provides evidence for a genetic mood disorders spectrum.

genomics

SplitMEM: Graphical pan-genome analysis with suffix skips

Motivation: With the rise of improved sequencing technologies, genomics is expanding from a single reference per species paradigm into a more comprehensive pan-genome approach with multiple individuals represented and analyzed together. One of the most sophisticated data structures for representing an entire population of genomes is a compressed de Bruijn graph. The graph structure can robustly represent simple SNPs to complex structural variations far beyond what can be done from linear sequences alone. As such there is a strong need to develop algorithms that can efficiently construct and analyze these graphs. Results: In this paper we explore the deep topological relationships between the suffix tree and the compressed de Bruijn graph. We introduce a novel O(n log n) time and space algorithm called splitMEM, that directly constructs the compressed de Bruijn graph for a pan-genome of total length n. To achieve this time complexity, we augment the suffix tree with suffix skips, a new construct that allows us to traverse several suffix links in constant time, and use them to efficiently decompose maximal exact matches (MEMs) into the graph nodes. We demonstrate the utility of splitMEM by analyzing the pan- genomes of 9 strains of Bacillus anthracis and 9 strains of Escherichia coli to reveal the properties of their core genomes. Availability: The source code and documentation are available open- source at http://splitmem.sourceforge.net

Bioinformatics

Assembling Large Genomes with Single-Molecule Sequencing and Locality Sensitive Hashing

We report reference-grade de novo assemblies of four model organisms and the human genome from single-molecule, real-time (SMRT) sequencing. Long-read SMRT sequencing is routinely used to finish microbial genomes, but the available assembly methods have not scaled well to larger genomes. Here we introduce the MinHash Alignment Process (MHAP) for efficient overlapping of noisy, long reads using probabilistic, locality-sensitive hashing. Together with Celera Assembler, MHAP was used to reconstruct the genomes of Escherichia coli, Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster, and human from high-coverage SMRT sequencing. The resulting assemblies include fully resolved chromosome arms and close persistent gaps in these important reference genomes, including heterochromatic and telomeric transition sequences. For D. melanogaster, MHAP achieved a 600-fold speedup relative to prior methods and a cloud computing cost of a few hundred dollars. These results demonstrate that single-molecule sequencing alone can produce near-complete eukaryotic genomes at modest cost.

Bioinformatics

The rate and molecular spectrum of spontaneous mutations in the GC-rich multi-chromosome genome of Burkholderia cenocepacia

Spontaneous mutations are ultimately essential for evolutionary change and are also the root cause of many diseases. However, until recently, both biological and technical barriers have prevented detailed analyses of mutation profiles, constraining our understanding of the mutation process to a few model organisms and leaving major gaps in our understanding of the role of genome content and structure on mutation. Here, we present a genome-wide view of the molecular mutation spectrum in Burkholderia cenocepacia, a clinically relevant pathogen with high %GC-content and multiple chromosomes. We find that B. cenocepacia has low genome-wide mutation rates with insertion-deletion mutations biased towards deletions, consistent with the idea that deletion pressure reduces prokaryotic genome sizes. Unlike prior studies of other organisms, mutations in B. cenocepacia are not AT-biased, which suggests that at least some genomes with high %GC-content experience unusual base-substitution mutation pressure. Importantly, we also observe variation in both the rates and spectra of mutations among chromosomes and elevated G:C>T:A transversions in late-replicating regions. Thus, although some patterns of mutation appear to be highly conserved across cellular life, others vary between species and even between chromosomes of the same species, potentially influencing the evolution of nucleotide composition and genome architecture.

Evolutionary Biology

Natural Selection Shapes the Mosaic Ancestry of the Drosophila Genetic Reference Panel and the D. melanogaster Reference Genome

North American populations of Drosophila melanogaster are thought to derive from both European and African source populations, but despite their importance for genetic research, patterns of admixture along their genomes are essentially undocumented. Here, I infer geographic ancestry along genomes of the Drosophila Genetic Reference Panel (DGRP) and the D. melanogaster reference genome. Overall, the proportion of African ancestry was estimated to be 20% for the DGRP and 9% for the reference genome. Based on the size of admixture tracts and the approximate timing of admixture, I estimate that the DGRP population underwent roughly 13.9 generations per year. Notably, ancestry levels varied strikingly among genomic regions, with significantly less African introgression on the X chromosome, in regions of high recombination, and at genes involved in specific processes such as circadian rhythm. An important role for natural selection during the admixture process was further supported by a genome-wide signal of ancestry disequilibrium, in that many between-chromosome pairs of loci showed a deficiency of Africa-Europe allele combinations. These results support the hypothesis that admixture between partially genetically isolated Drosophila populations led to natural selection against incompatible genetic variants, and that this process is ongoing. The ancestry blocks inferred here may be relevant for the performance of reference alignment in this species, and may bolster the design and interpretation of many population genetic and association mapping studies.

Evolutionary Biology

LINKS: Scaffolding genome assemblies with kilobase-long nanopore reads

MotivationOwing to the complexity of the assembly problem, we do not yet have complete genome sequences. The difficulty in assembling reads into finished genomes is exacerbated by sequence repeats and the inability of short reads to capture sufficient genomic information to resolve those problematic regions. Established and emerging long read technologies show great promise in this regard, but their current associated higher error rates typically require computational base correction and/or additional bioinformatics preprocessing before they could be of value. We present LINKS, the Long Interval Nucleotide K-mer Scaffolder algorithm, a solution that makes use of the information in error-rich long reads, without the need for read alignment or base correction. We show how the contiguity of an ABySS E. coli K-12 genome assembly could be increased over five-fold by the use of beta-released Oxford Nanopore Ltd. (ONT) long reads and how LINKS leverages long-range information in S. cerevisiae W303 ONT reads to yield an assembly with less than half the errors of competing applications. Re-scaffolding the colossal white spruce assembly draft (PG29, 20 Gbp) and how LINKS scales to larger genomes is also presented. We expect LINKS to have broad utility in harnessing the potential of long reads in connecting high-quality sequences of small and large genome assembly drafts.\n\nAvailabilityhttp://www.bcgsc.ca/bioinformatics/software/links\n\nContactrwarren@bcgsc.ca

Bioinformatics

Entire genome transcription across evolutionary time exposes non-coding DNA to de novo gene emergence

Even in the best studied Mammalian genomes, less than 5% of the total genome length is annotated as exonic. However, deep sequencing analysis in humans has shown that around 40% of the genome may be covered by poly-adenylated non-coding transcripts occurring at low levels1. Their functional significance is unclear2,3, and there has been a dispute whether they should be considered as noise of the transcriptional machinery4,5. We propose that if such transcripts show some evolutionary stability they will serve as substrates for de novo gene evolution, i.e. gene emergence out of non-coding DNA6-8. Here, we characterize the phylogenetic turnover of low-level poly-adenylated transcripts in a comprehensive sampling of populations, sub-species and species of the genus Mus, spanning a phylogenetic distance of about 10 Myr. We find evidence for more evolutionary stable gains of transcription than losses among closely related taxa, balanced by a loss of older transcripts across the whole phylogeny. We show that adding taxa increases the genomic transcript coverage and that no major transcript-free islands exist over time. This suggests that the entire genome can be transcribed into polyadenylated RNA when viewed at an evolutionary time scale. Thus, any part of the \"non-coding\" genome can become subject to evolutionary functionalization via de novo gene evolution.

Evolutionary Biology

A Statistical Framework to Predict Functional Non-Coding Regions in the Human Genome Through Integrated Analysis of Annotation Data

Identifying functional regions in the human genome is a major goal in human genetics. Great efforts have been made to functionally annotate the human genome either through computational predictions, such as genomic conservation, or high-throughput experiments, such as the ENCODE project. These efforts have resulted in a rich collection of functional annotation data of diverse types that need to be jointly analyzed for integrated interpretation and annotation. Here we present GenoCanyon, a whole-genome annotation method that performs unsupervised statistical learning using 22 computational and experimental annotations thereby inferring the functional potential of each position in the human genome. With GenoCanyon, we are able to predict many of the known functional regions. The ability of predicting functional regions as well as its generalizable statistical framework makes GenoCanyon a unique and powerful tool for whole-genome annotation. The GenoCanyon web server is available at http://genocanyon.med.yale.edu

Bioinformatics

Genome Wide Variant Analysis of Simplex Autism Families with an Integrative Clinical-Bioinformatics Pipeline

Autism spectrum disorders (ASD) are a group of developmental disabilities that affect social interaction, communication and are characterized by repetitive behaviors. There is now a large body of evidence that suggests a complex role of genetics in ASD, in which many different loci are involved. Although many current population scale genomic studies have been demonstrably fruitful, these studies generally focus on analyzing a limited part of the genome or use a limited set of bioinformatics tools. These limitations preclude the analysis of genome-wide perturbations that may contribute to the development and severity of ASD-related phenotypes. To overcome these limitations, we have developed and utilized an integrative clinical and bioinformatics pipeline for generating a more complete and reliable set of genomic variants for downstream analyses. Our study focuses on the analysis of three simplex autism families consisting of one affected child, unaffected parents, and one unaffected sibling. All members were clinically evaluated and widely phenotyped. Genotyping arrays and whole genome sequencing were performed on each member, and the resulting sequencing data were analyzed using a variety of available bioinformatics tools. We searched for rare variants of putative functional impact that were found to be segregating according to de-novo, autosomal recessive, x-linked, mitochondrial and compound heterozygote transmission models. The resulting candidate variants included three small heterozygous CNVs, a rare heterozygous de novo nonsense mutation in MYBBP1A located within exon 1, and a novel de novo missense variant in LAMB3. Our work demonstrates how more comprehensive analyses that include rich clinical data and whole genome sequencing data can generate reliable results for use in downstream investigations. We are moving to implement our framework for the analysis and study of larger cohorts of families, where statistical rigor can accompany genetic findings.

Genetics

Core variability in substitution rates and the basal sequence characteristics of the human genome

Accurate knowledge on the core components of substitution rates is of vital importance to understand genome evolution and dynamics. By performing a single-genome and direct analysis of 39894 retrotransposon remnants, we reveal germline sequence-dependent nucleotide substitution rates that can be assigned to each position in the human genome. Benefiting from the data made available in such detail, we show that a simulated genome, generated by equilibrating a random DNA sequence solely using our rate constants, exhibits nucleotide organisation close to that in the human genome. We next generate the germline basal substitution propensity (BSP) profile of the human genome and show a decreased tendency of moieties with low BSP to undergo somatic mutations in many cancer types.

Evolutionary Biology

Natural selection and recombination rate variation shape nucleotide polymorphism across the genomes of three related Populus species.

A central aim of evolutionary genomics is to identify the relative roles that various evolutionary forces have played in generating and shaping genetic variation within and among species. Here we use whole-genome re-sequencing data to characterize and compare genome-wide patterns of nucleotide polymorphism, site frequency spectrum and population-scaled recombination rates in three species of Populus: P. tremula, P. tremuloides and P. trichocarpa. We find that P. tremuloides has the highest level of genome-wide variation, skewed allele frequencies and population-scaled recombination rates, whereas P. trichocarpa harbors the lowest. Our findings highlight multiple lines of evidence suggesting that natural selection, both due to purifying and positive selection, has widely shaped patterns of nucleotide polymorphism at linked neutral sites in all three species. Differences in effective population sizes and rates of recombination are largely explaining the disparate magnitudes and signatures of linked selection we observe among species. The present work provides the first phylogenetic comparative study at genome-wide scale in forest trees. This information will also improve our ability to understand how various evolutionary forces have interacted to influence genome evolution among related species.

Evolutionary Biology

Evolutionary assembly patterns of prokaryotic genomes

Evolutionary innovation must occur in the context of some genomic background, which limits available evolutionary paths. For example, protein evolution by sequence substitution is constrained by epistasis between residues. In prokaryotes, evolutionary innovation frequently happens by macrogenomic events such as horizontal gene transfer (HGT). Previous work has suggested that HGT can be influenced by ancestral genomic content, yet the extent of such gene-level constraints has not yet been systematically characterized. Here, we evaluated the evolutionary impact of such constraints in prokaryotes, using probabilistic ancestral reconstructions from 634 extant prokaryotic genomes and a novel framework for detecting evolutionary constraints on HGT events. We identified 8,228 directional dependencies between genes, and demonstrated that many such dependencies reflect known functional relationships, including, for example, evolutionary dependencies of the photosynthetic enzyme RuBisCO. Modeling all dependencies as a network, we adapted an approach from graph theory to establish chronological precedence in the acquisition of different genomic functions. Specifically, we demonstrated that specific functions tend to be gained sequentially, suggesting that evolution in prokaryotes is governed by functional assembly patterns. Finally, we showed that these dependencies are universal rather than clade-specific and are often sufficient for predicting whether or not a given ancestral genome will acquire specific genes. Combined, our results indicate that evolutionary innovation via HGT is profoundly constrained by epistasis and historical contingency, similar to the evolution of proteins and phenotypic characters, and suggest that the emergence of specific metabolic and pathological phenotypes in prokaryotes can be predictable from current genomes.

Evolutionary Biology

Detecting Heterogeneity in Population Structure Across the Genome in Admixed Populations

The genetic structure of human populations is often characterized by aggregating measures of ancestry across the autosomal chromosomes. While it may be reasonable to assume that population structure patterns are similar genome-wide in relatively homogeneous populations, this assumption may not be appropriate for admixed populations, such as Hispanics and African Americans, with recent ancestry from two or more continents. Recent studies have suggested that systematic ancestry differences can arise at genomic locations in admixed populations as a result of selection and non-random mating. Here, we propose a method, which we refer to as the chromosomal ancestry differences (CAnD) test, for detecting heterogeneity in population structure across the genome. CAnD uses local ancestry inferred from SNP genotype data to identify chromosomes harboring genomic regions with ancestry contributions that are significantly different than expected. In simulation studies with real genotype data from Phase III of the HapMap Project, we demonstrate the validity and power of CAnD. We apply CAnD to the HapMap Mexican American (MXL) and African American (ASW) population samples; in this analysis the software RFMix is used to infer local ancestry at genomic regions assuming admixing from Europeans, West Africans, and Native Americans. The CAnD test provides strong evidence of heterogeneity in population structure across the genome in the MXL sample (p = 4e - 05), which is largely driven by elevated Native American ancestry and deficit of European ancestry on the X chromosomes. Among the ASW, all chromosomes are largely African derived and no heterogeneity in population structure is detected in this sample.

Genetics