Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Complete avian malaria parasite genomes reveal host-specific parasite evolution in birds and mammals

Avian malaria parasites are prevalent around the world, and infect a wide diversity of bird species. Here we report the sequencing and analysis of high quality draft genome sequences for two avian malaria species, Plasmodium relictum and Plasmodium gallinaceum. We identify 50 genes that are specific to avian malaria, located in an otherwise conserved core of the genome that shares gene synteny with all other sequenced malaria genomes. Phylogenetic analysis suggests that the avian malaria species form an outgroup to the mammalian Plasmodium species and using amino acid divergence between species, we estimate the avian and mammalian-infective lineages diverged in the order of 10 million years ago. Consistent with their phylogenetic position, we identify orthologs of genes that had previously appeared to be restricted to the clades of parasites containing P. falciparum and P. vivax - the species with the greatest impact on human health. From these orthologs, we explore differential diversifying selection across the genus and show that the avian lineage is remarkable in the extent to which invasion related genes are evolving. The subtelomeres of the P. relictum and P. gallinaceum genomes contain several novel gene families, including an expanded surf multigene family. We also identify an expansion of reticulocyte binding protein homologs in P. relictum and within these proteins, we detect distinct regions that are specific to non-human primate, humans, rodent and avian hosts. For the first time in the Plasmodium lineage we find evidence of transposable elements, including several hundred fragments of LTR-retrotransposons in both species and an apparently complete LTR-retrotransposon in the genome of P. gallinaceum.

genomics

An annotated draft genome for Radix auricularia(Gastropoda, Mollusca)

Molluscs are the second most species-rich phylum in the animal kingdom, yet only eleven genomes of this group have been published so far. Here, we present the draft genome sequence of the pulmonate freshwater snail Radix auricularia. Six whole genome shotgun libraries with different layouts were sequenced. The resulting assembly comprises 4,823 scaffolds with a cumulative length of 910 Mb and an overall read coverage of 72x. The assembly contains 94.6 % of a metazoan core gene collection, indicating an almost complete coverage of the coding fraction. The discrepancy of ~690 Mb compared to the estimated genome size of R. auricularia (1.6 Gb) results from a high repeat content of 70 % mainly comprising DNA transposons. The annotation of 17,338 protein coding genes was supported by the use of publicly-available transcriptome data. This draft will serve as starting point for further genomic and population genetic research in this scientifically important phylum.

genomics

Recombinant DNA resources for the comparative genomics of Ancylostoma ceylanicum.

We describe the construction and initial characterization of genomic resources (a set of recombinant DNA libraries, representing in total over 90,000 independent plasmid clones), originating from the genome of a hamster adapted hookworm, Ancylostoma ceylanicum. First, with the improved methodology, we generated sets of SL1 (5 -linker - GGTTAATTACCCAAGTTTGAG), and captured cDNAs from two different hookworm developmental stages: pre-infective L3 and parasitic adults. Second, we constructed a small insert (2-10kb) genomic library. Third, we generated a Bacterial Artificial Chromosome library (30-60kb). To evaluate the quality of our libraries we characterized sequence tags on randomly chosen clones and with first pass screening we generated almost a hundred novel hookworm sequence tags. The sequence tags detected two broad classes of genes: i. conserved nematode genes and ii. putative hookworm-specific proteins. Importantly, some of the identified genes encode proteins of general interest including potential targets for hookworm control. Additionally, we identified a syntenic region in the mitochondrial genome, where the gene order is shared between the free-living nematode C. elegans and A. ceylanicum. Our results validate the use of recombinant DNA resources for comparative genomics of nematodes, including the free-living genetic model organism C. elegans and closely related parasitic species. We discuss the potential and relevance of Ancylostoma ceylanicum data and resources generated by the recombinant DNA approach.

genomics

10 Simple Rules for Sharing Human Genomic Data

Introduction Introduction Conclusion Competing Interests References Delivery of the promise of precision medicine relies heavily on human genomic data sharing. Sharing genome data generated through publicly funded projects maximises return on investment from taxpayer funds and increases the likelihood of obtaining funding in future rounds [1]. More importantly, genome data sharing makes it possible for other scientists to reuse existing datasets for further research and constitutes a direct measure of the current advancement in risk prediction, diagnosis, and treatment for genomic disorders [2].\n\nSharing of human genomic data carries responsibilities to protect confidentiality and the privacy of research participants [3]. In certain cases data sharing may be complicated or limited by agreements with ...

genomics

Genome-wide analysis of ivermectin response by Onchocerca volvulus reveals that genetic drift and soft selective sweeps contribute to loss of drug sensitivity

BackgroundTreatment of onchocerciasis using mass ivermectin administration has reduced morbidity and transmission throughout Africa and Central/South America. Mass drug administration is likely to exert selection pressure on parasites, and phenotypic and genetic changes in several Onchocerca volvulus populations from Cameroon and Ghana - exposed to more than a decade of regular ivermectin treatment - have raised concern that sub-optimal responses to ivermectins anti-fecundity effect are becoming more frequent and may spread.\n\nMethodology/Principal FindingsPooled next generation sequencing (Pool-seq) was used to characterise genetic diversity within and between 108 adult female worms differing in ivermectin treatment history and response. Genome-wide analyses revealed genetic variation that significantly differentiated good responder (GR) and sub-optimal responder (SOR) parasites. These variants were not randomly distributed but clustered in ~31 quantitative trait loci (QTLs), with little overlap in putative QTL position and gene content between countries. Published candidate ivermectin SOR genes were largely absent in these regions; QTLs differentiating GR and SOR worms were enriched for genes in molecular pathways associated with neurotransmission, development, and stress responses. Finally, single worm genotyping demonstrated that geographic isolation and genetic change over time (in the presence of drug exposure) had a significantly greater role in shaping genetic diversity than the evolution of SOR.\n\nConclusions/SignificanceThis study is one of the first genome-wide association analyses in a parasitic nematode, and provides insight into the genomics of ivermectin response and population structure of O. volvulus. We argue that ivermectin response is a polygenically-determined quantitative trait in which identical or related molecular pathways but not necessarily individual genes likely determine the extent of ivermectin response in different parasite populations. Furthermore, we propose that genetic drift rather than genetic selection of SOR is the underlying driver of population differentiation, which has significant implications for the emergence and potential spread of SOR within and between these parasite populations.\n\nAuthor summaryOnchocerciasis is a human parasitic disease endemic across large areas of Sub-Saharan Africa, where more that 99% of the estimated 100 million people globally at-risk live. The microfilarial stage of Onchocerca volvulus causes pathologies ranging from mild itching to visual impairment and ultimately, irreversible blindness. Mass administration of ivermectin kills microfilariae and has an anti-fecundity effect on adult worms by temporarily inhibiting the development in utero and/or release into the skin of new microfilariae, thereby reducing morbidity and transmission. Phenotypic and genetic changes in some parasite populations that have undergone multiple ivermectin treatments in Cameroon and Ghana have raised concern that sub-optimal response to ivermectins anti-fecundity effect may increase in frequency, reducing the impact of ivermectin-based control measures. We used next generation sequencing of small pools of parasites to define genome-wide genetic differences between phenotypically characterised good and sub-optimal responder parasites from Cameroon and Ghana, and identified multiple genomic regions differentiating the response types. These regions were largely different between parasites from both countries but revealed common molecular pathways that might be involved in determining the extent of response to ivermectins anti-fecundity effect. These data reveal a more complex than previously described pattern of genetic diversity among O. volvulus populations that differ in their geography and response to ivermectin treatment.

genomics

Network based conditional genome wide association analysis of human metabolomics

BackgroundGenome-wide association studies (GWAS) have identified hundreds of loci influencing complex human traits, however, their biological mechanism of action remains mostly unknown. Recent accumulation of functional genomics ( omics) including metabolomics data opens up opportunities to provide a new insight into the functional role of specific changes in the genome. Functional genomic data are characterized by high dimensionality, presence of (strong) statistical dependencies between traits, and, potentially, complex genetic control. Therefore, analysis of such data asks for development of specific statistical genetic methods.\n\nResultsWe propose a network-based, conditional approach to evaluate the impact of genetic variants on omics phenotypes (conditional GWAS, cGWAS). For each trait of interest, based on biological network, we select a set of other traits to be used as covariates in GWAS. The network could be reconstructed either from biological pathway databases or directly from the data. We evaluated our approach using data from a population-based KORA study (n=1,784, 1.7 M SNPs) with measured metabolomics data (151 metabolites) and demonstrated that our approach allows for identification of up to five additional loci not detected by conventional GWAS. We show that this gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.\n\nConclusionsWe demonstrate that in context of metabolomics network-based, conditional genome-wide association analysis is able to dramatically increase power of identification of loci with specific contra-intuitive pleiotropic architecture. Our method has modest computational costs, can utilize summary level GWAS data, and is applicable to other omics data types. We anticipate that application of our method to new and existing data sets will facilitate progress in understanding genetic bases of control of molecular and complex phenotypes.\n\nShort abstractWe propose a network-based, conditional approach for genome-wide analysis of multivariate omics phenotypes. Our methods can incorporate prior biological knowledge about biological pathways from external sources. We evaluated our approach using metabolomics data and demonstrated that our approach has bigger power and allows for identification of additional loci. We show that gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.

genomics

Multiplex PCR method for MinION and Illumina sequencing of Zika and other virus genomes directly from clinical samples

Genome sequencing has become a powerful tool for studying emerging infectious diseases; however, genome sequencing directly from clinical samples without isolation remains challenging for viruses such as Zika, where metagenomic sequencing methods may generate insufficient numbers of viral reads. Here we present a protocol for generating coding-sequence complete genomes comprising an online primer design tool, a novel multiplex PCR enrichment protocol, optimised library preparation methods for the portable MinION sequencer (Oxford Nanopore Technologies) and the Illumina range of instruments, and a bioinformatics pipeline for generating consensus sequences. The MinION protocol does not require an internet connection for analysis, making it suitable for field applications with limited connectivity. Our method relies on multiplex PCR for targeted enrichment of viral genomes from samples containing as few as 50 genome copies per reaction. Viral consensus sequences can be achieved starting with clinical samples in 1-2 days following a simple laboratory workflow. This method has been successfully used by several groups studying Zika virus evolution and is facilitating an understanding of the spread of the virus in the Americas.

genomics

A High Quality Assembly of the Nile Tilapia (Oreochromis niloticus) Genome Reveals the Structure of Two Sex Determination Regions

We report a high-quality assembly of the tilapia genome, a perciform fish important in aquaculture around the world. A homozygous clonal XX female Nile tilapia (Oreochromis niloticus) was sequenced to 44X coverage using Pacific Biosciences (PacBio) SMRT sequencing. Dozens of candidate de novo assemblies were generated and an optimal assembly (contig NG50 of 3.3Mbp) was selected using principal component analysis of likelihood scores calculated from several paired-end sequencing libraries. Comparison of the new assembly to the previous O. niloticus genome assembly reveals that recently duplicated portions of the genome are now well represented. The overall number genes in the new assembly increased by 27.3%, including a 67% increase in pseudogenes. The new tilapia genome assembly correctly represents two recent vasa gene duplication events that have been verified with BAC sequencing. At total of 146Mbp of additional transposable element sequence are now assembled, a large proportion of which are recent insertions. Large centromeric satellite repeats are assembled and annotated in cichlid fish for the first time. Finally, the new assembly identifies the long-range structure of both an ~9Mbp XY sex-determination region on LG1 in O. niloticus, and a ~50Mbp WZ sex-determination region on LG3 in the related species O. aureus. This study highlights the use of long read sequencing to correctly assemble recent duplications and to characterize repeat-filled regions of the genome.

genomics

Rapid de novo assembly of the European eel genome from nanopore sequencing reads

We have sequenced the genome of the endangered European eel using the MinION by Oxford Nanopore, and assembled these data using a novel algorithm specifically designed for large eukaryotic genomes. For this 860 Mbp genome, the entire computational process takes two days on a single CPU. The resulting genome assembly significantly improves on a previous draft based on short reads only, both in terms of contiguity (N50 1.2 Mbp) and structural quality. This combination of affordable nanopore sequencing and light-weight assembly promises to make high-quality genomic resources accessible for many non-model plants and animals.

genomics

Cas9-Assisted Targeting of CHromosome segments (CATCH) for targeted nanopore sequencing and optical genome mapping

Variations in the genetic code, from single point mutations to large structural or copy number alterations, influence susceptibility, onset, and progression of genetic diseases and tumor transformation. Next-generation sequencing analyses are unable to reliably capture aberrations larger than the typical sequencing read length of several hundred bases. Long-read, single-molecule sequencing methods such as SMRT and nanopore sequencing can address larger variations, but require costly whole genome analysis. Here we describe a method for isolation and enrichment of a large genomic region of interest for targeted analysis based on Cas9 excision of two sites flanking the target region and isolation of the excised DNA segment by pulsed field gel electrophoresis. The isolated target remains intact and is ideally suited for optical genome mapping and long-read sequencing at high coverage. In addition, analysis is performed directly on native genomic DNA that retains genetic and epigenetic composition without amplification bias. This method enables detection of mutations and structural variants as well as detailed analysis by generation of hybrid scaffolds composed of optical maps and sequencing data at a fraction of the cost of whole genome sequencing.

genomics

Signatures of long-term balancing selection in human genomes

Balancing selection maintains advantageous diversity in populations through various mechanisms. While extensively explored from a theoretical perspective, an empirical understanding of its prevalence and targets lags behind our knowledge of positive selection. Here we describe the Non-Central Deviation (NCD), a simple yet powerful statistic to detect long-term balancing selection (LTBS) that quantifies how close frequencies are to expectations under LTBS, and provides the basis for a neutrality test. NCD can be applied to a single locus or genomic data, and can be implemented considering only polymorphisms (NCD1) or also considering fixed differences with respect to an outgroup (NCD2) species. Incorporating fixed differences improves power, and NCD2 has higher power to detect LTBS in humans under different frequencies of the balanced allele(s) than other available methods. Applied to genome-wide data from African and European human populations, in both cases using chimpanzee as an outgroup, NCD2 shows that, albeit not prevalent, LTBS affects a sizable portion of the genome: about 0.6% of analyzed genomic windows and 0.8% of analyzed positions. Significant windows (p < 0.0001) contain 1.6% of SNPs in the genome, which disproportionally fall within exons and change protein sequence, but are not enriched in putatively regulatory sites. These windows overlap about 8% of the protein-coding genes, and these have larger number of transcripts than expected by chance even after controlling for gene length. Our catalog includes known targets of LTBS but a majority of them (90%) are novel. As expected, immune-related genes are among those with the strongest signatures, although most candidates are involved in other biological functions, suggesting that LTBS potentially influences diverse human phenotypes.

genomics

The Nuclear And Mitochondrial Genomes Of The Facultatively Eusocial Orchid Bee Euglossa dilemma

Bees provide indispensable pollination services to both agricultural crops and wild plant populations, and several species of bees have become important models for the study of learning and memory, plant-insect interactions and social behavior. Orchid bees (Apidae: Euglossini) are especially important to the fields of pollination ecology, evolution, and species conservation. Here we report the nuclear and mitochondrial genome sequences of the orchid bee Euglossa dilemma Bembe & Eltz. Euglossa dilemma was selected because it is widely distributed, highly abundant, and it was recently naturalized in the southeastern United States. We provide a high-quality assembly of the 3.3 giga-base genome, and an official gene set of 15,904 gene annotations. We find high conservation of gene synteny with the honey bee throughout 80 million years of divergence time. This genomic resource represents the first draft genome of the orchid bee genus Euglossa, and the first draft orchid bee mitochondrial genome, thus representing a valuable resource to the research community.

genomics

Nanopore sequencing and assembly of a human genome with ultra-long reads

Nanopore sequencing is a promising technique for genome sequencing due to its portability, ability to sequence long reads from single molecules, and to simultaneously assay DNA methylation. However until recently nanopore sequencing has been mainly applied to small genomes, due to the limited output attainable. We present nanopore sequencing and assembly of the GM12878 Utah/Ceph human reference genome generated using the Oxford Nanopore MinION and R9.4 version chemistry. We generated 91.2 Gb of sequence data ([~]30x theoretical coverage) from 39 flowcells. De novo assembly yielded a highly complete and contiguous assembly (NG50 [~]3Mb). We observed considerable variability in homopolymeric tract resolution between different basecallers. The data permitted sensitive detection of both large structural variants and epigenetic modifications. Further we developed a new approach exploiting the long-read capability of this system and found that adding an additional 5x-coverage of ultra-long reads (read N50 of 99.7kb) more than doubled the assembly contiguity. Modelling the repeat structure of the human genome predicts extraordinarily contiguous assemblies may be possible using nanopore reads alone. Portable de novo sequencing of human genomes may be important for rapid point-of-care diagnosis of rare genetic diseases and cancer, and monitoring of cancer progression. The complete dataset including raw signal is available as an Amazon Web Services Open Dataset at: https://github.com/nanopore-wgs-consortium/NA12878.

genomics

Reconstructing The Gigabase Plant Genome Of Solanum pennellii Using Nanopore Sequencing

Recent updates in sequencing technology have made it possible to obtain Gigabases of sequence data from one single flowcell. Prior to this update, the nanopore sequencing technology was mainly used to analyze and assemble microbial samples1-3. Here, we describe the generation of a comprehensive nanopore sequencing dataset with a median fragment size of 11,979 bp for the wild tomato species Solanum pennellii featuring an estimated genome size of ca 1.0 to 1.1 Gbases. We describe its genome assembly to a contig N50 of 2.5 MB using a pipeline comprising a Canu4 pre-processing and a subsequent assembly using SMARTdenovo. We show that the obtained nanopore based de novo genome reconstruction is structurally highly similar to that of the reference S. pennellii LA7165 genome but has a high error rate caused mostly by deletions in homopolymers. After polishing the assembly with Illumina short read data we obtained an error rate of <0.02 % when assessed versus the same Illumina data. More importantly however we obtained a gene completeness of 96.53% which even slightly surpasses that of the reference S. pennellii genome5. Taken together our data indicate such long read sequencing data can be used to affordably sequence and assemble Gbase sized diploid plant genomes.\n\nRaw data is available at http://www.plabipd.de/portal/solanum-pennellii and has been deposited as PRJEB19787.

genomics

A Duplication Lost In Sugarcane Hybrids Revealed By Chloroplast Genome Assembly Of Wild Species Saccharum officinarum

Sugarcane is a crop of paramount importance for sustainable energy. Modern sugarcane cultivars are derived from interspecific crosses between the two wild species Saccharum officinarum and Saccharum spontaneum and this event occurred very early in the sugarcane domestication history. This hybridization allowed the generation of cultivars with complex aneuploidy genomes containing 100-130 chromosomes that are unequally inherited - ~80% from S. officinarum, ~10% from S. spontaneum and ~10% from inter-specific crosses. Several studies have highlighted the importance of chloroplast genomes (cpDNA) to investigate hybridization events in plant lineages. Few sugarcane cpDNAs have been assembled and published, including those from sugarcane hybrids. However, cpDNAs of wild Saccharum species remains unexplored. In the present study, we used whole-genome sequencing data to survey the chloroplast genome of the wild sugarcane species S. officinarum. Illumina sequencing technology was used for assembly 142,234 bp of S.officinarum cpDNA with 2,065,893 reads and 1043x of coverage. The analysis of the S. officinarum cpDNA revealed a notable difference in the LSC region of wild and cultivated sugarcanes. Chloroplasts of sugarcane cultivars showed a loss of a duplicated fragment with 1,031 bp in the beginning of the LSC region, which decreased the chloroplast gene content in hybrids. Based on these results, we propose the comparative analysis of organelle genomes as a very important tool for deciphering and understanding hybrid Saccharum lineages.

genomics

Dissecting the Causal Mechanism of X-Linked Dystonia-Parkinsonism by Integrating Genome and Transcriptome Assembly

X-linked Dystonia-Parkinsonism (XDP) is a Mendelian neurodegenerative disease endemic to the Philippines. We integrated genome and transcriptome assembly with induced pluripotent stem cell-based modeling to identify the XDP causal locus and potential pathogenic mechanism. Genome sequencing identified novel variation that was shared by all probands and three recombination events that narrowed the causal locus to a genomic segment including TAF1. Transcriptome assembly in neural derivative cells discovered novel TAF1 transcripts, including a truncated transcript exclusively observed in probands that involved aberrant splicing and intron retention (IR) associated with a SINE-VNTR-Alu (SVA)-type retrotransposon insertion. This IR correlated with decreased expression of the predominant TAF1 transcript and altered expression of neurodevelopmental genes; both the IR and aberrant TAF1 expression patterns were rescued by CRISPR/Cas9 excision of the SVA. These data suggest a unique genomic cause of XDP and may provide a roadmap for integrative genomic studies in other unsolved Mendelian disorders.\n\nHighlights O_LIGenome assembly narrows the XDP causal locus to a segment including TAF1\nC_LIO_LIXDP-specific SVA insertion induces intron retention and down-regulation of TAF1\nC_LIO_LICRISPR/Cas9 excision of SVA rescues aberrant splicing and cTAF1 expression in XDP\nC_LIO_LIGene networks perturbed in proband cells associate to synapse and neurodevelopment\nC_LI

genomics

An integrative strategy to identify the entire protein coding potential of prokaryotic genomes by proteogenomics

Accurate annotation of all protein-coding sequences (CDSs) is an essential prerequisite to fully exploit the rapidly growing repertoire of completely sequenced prokaryotic genomes. However, large discrepancies among the number of CDSs annotated by different resources, missed functional short open reading frames (sORFs), and overprediction of spurious ORFs represent serious limitations.\n\nOur strategy towards accurate and complete genome annotation consolidates CDSs from multiple reference annotation resources, ab initio gene prediction algorithms and in silico ORFs in an integrated proteogenomics database (iPtgxDB) that covers the entire protein-coding potential of a prokaryotic genome. By extending the PeptideClassifier concept of unambiguous peptides for prokaryotes, close to 95% of the identifiable peptides imply one distinct protein, largely simplifying downstream analysis. Searching a comprehensive Bartonella henselae proteomics dataset against such an iPtgxDB allowed us to unambiguously identify novel ORFs uniquely predicted by each resource, including lipoproteins, differentially expressed and membrane-localized proteins, novel start sites and wrongly annotated pseudogenes. Most novelties were confirmed by targeted, parallel reaction monitoring mass spectrometry, including unique ORFs and variants identified in a re-sequenced laboratory strain that are not present in its reference genome. We demonstrate the general applicability of our strategy for genomes with varying GC content and distinct taxonomic origin, and release iPtgxDBs for B. henselae, Bradyrhozibium diazoefficiens and Escherichia coli as well as the software to generate such proteogenomics search databases for any prokaryote.

genomics

The pomegranate (Punica granatum L.) genome provides insights into fruit quality and ovule developmental biology

Pomegranate (Punica granatum L.) with an uncertain taxonomic status has an ancient cultivation history, and has become an emerging fruit due to its attractive features such as the bright red appearance and the high abundance of medicinally valuable ellagitannin-based compounds in its peel and aril. However, the absence of genomic resources has restricted further elucidating genetics and evolution of these interesting traits. Here we report a 274-Mb high-quality draft pomegranate genome sequence, which covers approximately 81.5% of the estimated 336 Mb genome, consists of 2,177 scaffolds with an N50 size of 1.7 Mb, and contains 30,903 genes. Phylogenomic analysis supported that pomegranate belongs to the Lythraceae family rather than the monogeneric Punicaceae family, and comparative analyses showed that pomegranate and Eucalyptus grandis shares the paleotetraploidy event. Integrated genomic and transcriptomic analyses provided insights into the molecular mechanisms underlying the biosynthesis of ellagitannin-based compounds, the color formation in both peels and arils during pomegranate fruit development, and the unique ovule development processes that are characteristic of pomegranate. This genome sequence represents the first reference in Lythraceae, providing an important resource to expand our understanding of some unique biological processes and to facilitate both comparative biology studies and crop breeding.

genomics