Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

An Improved Genome Assembly of Azadirachta indica A. Juss.

Neem (Azadirachta indica A. Juss.), an evergreen tree of the Meliaceae family, is known for its medicinal, cosmetic, pesticidal and insecticidal properties. We had previously sequenced and published the draft genome of the plant, using mainly short read sequencing data. In this report, we present an improved genome assembly generated using additional short reads from Illumina and long reads from Pacific Biosciences SMRT sequencer. We assembled short reads and error-corrected long reads using Platanus, an assembler designed to perform well for heterozygous genomes. The updated genome assembly (v2.0) yielded 3- and 3.5-fold increase in N50 and N75, respectively; 2.6-fold decrease in the total number of scaffolds; 1.25-fold increase in the number of valid transcriptome alignments; 13.4-fold less mis-assembly and 1.85-fold increase in the percentage repeat, over the earlier assembly (v1.0). The current assembly also maps better to the genes known to be involved in the terpenoid biosynthesis pathway. Together, the data represents an improved assembly of the A. indica genome.\n\nThe raw data described in this manuscript are submitted to the NCBI Short Read Archive under the accession numbers SRX1074131, SRX1074132, SRX1074133, and SRX1074134 (SRP013453).

Bioinformatics

Predictive computational phenotyping and biomarker discovery using reference-free genome comparisons

BackgroundThe identification of genomic biomarkers is a key step towards improving diagnostic tests and therapies. We present a reference-free method for this task that relies on a k-mer representation of genomes and a machine learning algorithm that produces intelligible models. The method is computationally scalable and well-suited for whole genome sequencing studies.\n\nResultsThe method was validated by generating models that predict the antibiotic resistance of C. difficile, M. tuberculosis, P. aeruginosa, and S. pneumoniae for 17 antibiotics. The obtained models are accurate, faithful to the biological pathways targeted by the antibiotics, and they provide insight into the process of resistance acquisition. Moreover, a theoretical analysis of the method revealed tight statistical guarantees on the accuracy of the obtained models, supporting its relevance for genomic biomarker discovery.\n\nConclusionsOur method allows the generation of accurate and interpretable predictive models of phenotypes, which rely on a small set of genomic variations. The method is not limited to predicting antibiotic resistance in bacteria and is applicable to a variety of organisms and phenotypes. Kover, an efficient implementation of our method, is open-source and should guide biological efforts to understand a plethora of phenotypes (http://github.com/aldro61/kover/).

Bioinformatics

Genomic positional conservation identifies topological anchor point (tap)RNAs linked to developmental loci

The mammalian genome is transcribed into large numbers of long noncoding RNAs (lncRNAs), but the definition of functional lncRNA groups has proven difficult, partly due to their low sequence conservation and lack of identified shared properties. Here we consider positional conservation across mammalian genomes as an indicator of functional commonality. We identify 665 conserved lncRNA promoters in mouse and human genomes that are preserved in genomic position relative to orthologous coding genes. The identified positionally conserved lncRNA genes are primarily associated with developmental transcription factor loci with which they are co-expressed in a tissue-specific manner. Strikingly, over half of all positionally conserved RNAs in this set are linked to distinct chromatin organization structures, overlapping the binding sites for the CTCF chromatin organizer and located at chromatin loop anchor points and borders of topologically associating domains (TADs). These topological anchor point (tap)RNAs possess conserved sequence domains that are enriched in potential recognition motifs for Zinc Finger proteins. Characterization of these noncoding RNAs and their associated coding genes shows that they are functionally connected: they regulate each others expression and influence the metastatic phenotype of cancer cells in vitro in a similar fashion. Thus, interrogation of positionally conserved lncRNAs identifies a new subset of tapRNAs with shared functional properties. These results provide a large dataset of lncRNAs that conform to the \"extended gene\" model, in which conserved developmental genes are genomically and functionally linked to regulatory lncRNA loci across mammalian evolution.

Molecular Biology

trio-sga: facilitating de novo assembly of highly heterozygous genomes with parent-child trios

MotivationMost DNA sequence in diploid organisms is found in two copies, one contributed by the mother and the other by the father. The high density of differences between the maternally and paternally contributed sequences (heterozygous sites) in some organisms makes de novo genome assembly very challenging, even for algorithms specifically designed to deal with these cases. Therefore, various approaches, most commonly inbreeding in the laboratory, are used to reduce heterozygosity in genomic data prior to assembly. However, many species are not amenable to these techniques.\n\nResultsWe introduce trio-sga, a set of three algorithms designed to take advantage of mother-father-offspring trio sequencing to facilitate better quality genome assembly in organisms with moderate to high levels of heterozygosity. Two of the algorithms use haplotype phase information present in the trio data to eliminate the majority of heterozygous sites before the assembly commences. The third algorithm is designed to reduce sequencing costs by enabling the use of parents reads in the assembly of the genome of the offspring. We test these algorithms on a simulated trio from four hap-loid datasets, and further demonstrate their performance by assembling three highly heterozygous Heliconius butterfly genomes. While the implementation of trio-sga is tuned towards Illumina-generated data, we note that the trio approach to reducing heterozygosity is likely to have cross-platform utility for de novo assembly.

Bioinformatics

Going down the rabbit hole: a review on how to link genome-wide data with ecology and evolution in natural populations

O_LICharacterizing species history and identifying loci underlying local adaptation is crucial in functional ecology, evolutionary biology, conservation and agronomy. The ongoing and constant improvement of next-generation sequencing (NGS) techniques has facilitated the production of an ever-increasing number of genetic markers across genomes of non-model species.\nC_LIO_LIThe study of variation in these markers across natural populations has deepened the understanding of how population history and selection act on genomes. Population genomics now provides tools to better integrate selection into a historical framework, and take into account selection when reconstructing demographic history. However, this improvement has come with a burst of analytical tools that can confuse users.\nC_LIO_LISuch confusion can limit the amount of information effectively retrieved from complex genomic datasets. In addition, the lack of a unified analytical pipeline impairs the diffusion of the most recent analytical tools into fields like conservation biology.\nC_LIO_LITo address this need, we describe possible analytical protocols and link these with more than 70 methods dealing with genome-scale datasets. We summarise the strategies they use to infer demographic history and selection, and discuss some of their limitations. A website listing these methods is available at www.methodspopgen.com.\nC_LI

Evolutionary Biology

Scaffolding and Completing Genome Assemblies in Real-time with Nanopore Sequencing

Genome assemblies obtained from short read sequencing technologies are often fragmented into many contigs because of the abundance of repetitive sequences. Long read sequencing technologies allow the generation of reads spanning most repeat sequences, providing the opportunity to complete these genome assemblies. However, substantial amounts of sequence data and computational resources are required to overcome the high per-base error rate inherent to these technologies. Furthermore, most existing methods only assemble the genomes after sequencing has completed which could result in either generation of more sequence data at greater cost than required or a low-quality assembly if insufficient data are generated. Here we present the first computational method which utilises real-time nanopore sequencing to scaffold and complete short-read assemblies while the long read sequence data is being generated. The method reports the progress of completing the assembly in real-time so users can terminate the sequencing once an assembly of sufficient quality and completeness is obtained. We use our method to complete four bacterial genomes and one eukaryotic genome, and show that it is able to construct more complete and more accurate assemblies, and at the same time, requires less sequencing data and computational resources than existing pipelines. We also demonstrate that the method can facilitate real-time analyses of positional information such as identification of bacterial genes encoded in plasmids and pathogenicity islands.

Bioinformatics

Not just methods: User expertise explains the variability of outcomes of genome-wide studies

Genome scan approaches promise to map genomic regions involved in adaptation of individuals to their environment. Outcomes of genome scans have been shown to depend on several factors including the underlying demography, the adaptive scenario, and the software or method used. We took advantage of a pedagogical experiment carried out during a summer school to explore the effect of an unexplored source of variability, which is the degree of user expertise.Participants were asked to analyze three simulated data challenges with methods presented during the summer school. In addition to submitting lists, participants evaluated a priori their level of expertise. We measured the quality of each genome scan analysis by computing a score that depends on false discovery rate and statistical power. In an easy and a difficult challenge, less advanced participants obtained similar scores compared to advanced ones, demonstrating that participants with little background in genome scan methods were able to learn how to use complex software after short introductory tutorials. However, in a challenge ofintermediate difficulty, advanced participants obtained better scores. To explain the difference, we introduce a probabilistic model that shows that a larger variation in scores is expected for SNPs of intermediate difficulty of detection. We conclude that practitioners should develop their statistical and computational expertise to follow the development of complex methods. To encourage training, we release the website of the summer school where users can submit lists of candidate loci, which will be scored and compared to the scores obtained by previous users.

Ecology

mmgenome: a toolbox for reproducible genome extraction from metagenomes

SummaryRecovery of population genomes is becoming a standard analysis in metagenomics and a multitude of different approaches exists. However, the workflows are complex, requiring data generation, binning, validation and finishing to generate high quality population genome bins. In addition, several different approaches are often used on the same dataset as the optimal strategy to extract a specific population genome varies. Here we introduce mmgenome: a toolbox for reproducible genome extraction from metagenomes. At the core of mmgenome is an R package that facilitates effortless integration of different binning strategies by collecting information on scaffolds. Genome binning is facilitated through integrated tools that support effortless visualizations, validation and calculation of key statistics. Full reproducibility and transparency is obtained through Rmarkdown, whereby every step can be recreated.\n\nAvailability and implementationThe binning framework of mmge-nome is implemented in R. Wrapper scripts for data generation and finishing is written in Perl. The mmgenome toolbox and associated step-by-step guides are available at http://madsal-bertsen.github.io/mmgenome/.\n\nContactma@bio.aau.dk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Impact of Sample Type and DNA Isolation Procedure on Genomic Inference of Microbiome Composition

Explorations of complex microbiomes using genomics greatly enhance our understanding about their diversity, biogeography, and function. The isolation of DNA from microbiome specimens is a key prerequisite for such examinations, but challenges remain in obtaining sufficient DNA quantities required for certain sequencing approaches, achieving accurate genomic inference of microbiome composition, and facilitating comparability of findings across specimen types and sequencing projects. These aspects are particularly relevant for the genomics-based global surveillance of infectious agents and antimicrobial resistance from different reservoirs. Here, we compare in a stepwise approach a total of eight commercially available DNA extraction kits and 16 procedures based on these for three specimen types (human feces, pig feces, and hospital sewage). We assess DNA extraction using spike-in controls, and different types of beads for bead-beating facilitating cell lysis. We evaluate DNA concentration, purity, and stability, and microbial community composition using 16S rRNA gene sequencing and for selected samples using shotgun metagenomic sequencing. Our results suggest that inferred community composition was dependent on inherent specimen properties as well as DNA extraction method. We further show that bead-beating or enzymatic treatment can increase the extraction of DNA from Gram-positive bacteria. Final DNA quantities could be increased by isolating DNA from a larger volume of cell lysate compared to standard protocols. Based on this insight, we designed an improved DNA isolation procedure optimized for microbiome genomics that can be used for the three examined specimen types and potentially also for other biological specimens. A standard operating procedure is available from: https://dx.doi.org/10.6084/m9.figshare.3475406.\n\nIMPORTANCESequencing-based analyses of microbiomes may lead to a breakthrough in our understanding of the microbial world associate with humans, animals, and the environment. Such insight could further the development of innovative ecosystem management approaches for the protection of our natural resources, and the design of more effective and sustainable solutions to prevent and control infectious diseases. Genome sequence information is an organism- (pathogen-) independent language that can be used across sectors, space, and time. Harmonized standards, protocols, and workflows for sample processing and analysis can facilitate the generation of such actionable information. In this study, we assessed several procedures for the isolation of DNA for next-generation sequencing. Our study highlights several important aspects to consider in the design and conduction of sequence-based analysis of microbiomes. We provide a standard operating procedure for the isolation of DNA from a range of biological specimens particularly relevant in clinical diagnostics and epidemiology.

Microbiology

Genome size variation and species diversity in salamander families

Salamanders (Urodela) have among the largest vertebrate genomes, ranging in size from 10 to 120 pg. Although changes in genome size often occur randomly and in the absence of selection pressure, non-random patterns of genome size variation are evident among specific vertebrate lineages. Several reports suggest a relationship between species richness and genome size, but the exact nature of that relationship remains unclear both within and across different taxonomic groups. Here we report i) a negative relationship between haploid genome size (C-value) and species richness at the family taxonomic level in salamander clades; ii) a correlation of C-value and species richness with clade crown-age but not with diversification rates; iii) strong associations between C-value and either geographical area or climatic niche rate. Finally, we report a relationship between C-value diversity and species diversity at both the family and genus level clades in urodeles.

Evolutionary Biology

Genome Editing With Targeted Deaminases

Precise genetic modifications are essential for biomedical research and gene therapy. Yet, traditional homology-directed genome editing is limited by the requirements for DNA cleavage, donor DNA template and the endogenous DNA break-repair machinery. Here we present programmable cytidine deaminases that enable site-specific cytidine to thymidine (C-to-T) genomic edits without the need for DNA cleavage. Our targeted deaminases are efficient and specific in Escherichia coli, converting a genomic C-to-T with 13% efficiency and 95% accuracy. Edited cells do not harbor unintended genomic abnormalities. These novel enzymes also function in human cells, leading to a site-specific C-to-T transition in 2.5% of cells with reduced toxicity compared with zinc-finger nucleases. Targeted deaminases therefore represent a platform for safer and effective genome editing in prokaryotes and eukaryotes, especially in systems where DSBs are toxic, such as human stem cells and repetitive elements targeting.

Bioengineering

Population-genomic inference of the strength and timing of selection against gene flow

The interplay of divergent selection and gene flow is key to understanding how populations adapt to local environments and how new species form. Here, we use DNA polymorphism data and genome-wide variation in recombination rate to jointly infer the strength and timing of selection, as well as the baseline level of gene flow under various demographic scenarios. We model how divergent selection leads to a genome-wide negative correlation between recombination rate and genetic differentiation among populations. Our theory shows that the selection density, i.e. the selection coefficient per base pair, is a key parameter underlying this relationship. We then develop a procedure for parameter estimation that accounts for the confounding effect of background selection. Applying this method to two datasets from Mimulus guttatus, we infer a strong signal of adaptive divergence in the face of gene flow between populations growing on and off phytotoxic serpentine soils. However, the genome-wide intensity of this selection is not exceptional compared to what M. guttatus populations may typically experience when adapting to local conditions. We also find that selection against genome-wide introgression from the selfing sister species M. nasutus has acted to maintain a barrier between these two species over at least the last 250 ky. Our study provides a theoretical framework for linking genome-wide patterns of divergence and recombination with the underlying evolutionary mechanisms that drive this differentiation.

Evolutionary Biology

Approximate Likelihood Inference of Complex Population Histories and Recombination from Multiple Genomes

We introduce ABLE (Approximate Blockwise Likelihood Estimation), a novel composite likelihood framework based on a recently introduced summary of sequence variation: the blockwise site frequency spectrum (bSFS). This simulation-based framework uses the the frequencies of bSFS configurations to jointly model demographic history and recombination and is explicitly designed to make inference using multiple whole genomes or genome-wide multi-locus data (e.g. RADSeq) catering to the needs of researchers studying model or non-model organisms respectively. The flexible nature of our method further allows for arbitrarily complex population histories using unphased and unpolarized whole genome sequences. In silico experiments demonstrate accurate parameter estimates across a range of divergence models with increasing complexity, and as a proof of principle, we infer the demographic history of the two species of orangutan from multiple genome sequences (over 160 Mbp in length) from each species. Our results indicate that the two orangutan species split approximately 650-950 thousand years ago but experienced a pulse of secondary contact much more recently, most likely during a period of low sea-level South East Asia ([~]300,000 years ago). Unlike previous analyses we can reject a history of continuous gene flow and co-estimate genome-wide recombination. ABLE is available for download at https://github.com/champost/ABLE.

Evolutionary Biology

The PhytoClust Tool for Metabolic Gene Clusters Discovery in Plant Genomes

The existence of Metabolic Gene Clusters (MGCs) in plant genomes has recently raised increased interest. Thus far, MGCs were commonly identified for pathways of specialized metabolism, mostly those associated with terpene type products. For efficient identification of novel MGCs computational approaches are essential. Here we present PhytoClust; a tool for the detection of candidate MGCs in plant genomes. The algorithm employs a collection of enzyme families related to plant specialized metabolism, translated into hidden Markov models, to mine given genome sequences for physically co-localized metabolic enzymes. Our tool accurately identifies previously characterized plant MBCs. An exhaustive search of 31 plant genomes detected 1232 and 5531 putative gene cluster types and candidates, respectively. Clustering analysis of putative MGCs types by species reflected plant taxonomy. Furthermore, enrichment analysis revealed taxa- and species-specific enrichment of certain enzyme families in MGCs. When operating through our web-interface, PhytoClust users can mine a genome either based on a list of known cluster types or by defining new cluster rules. Moreover, for selected plant species, the output can be complemented by co-expression analysis. Altogether, we envisage PhytoClust to enhance novel MGCs discovery which will in turn impact the exploration of plant metabolism.

Plant Biology

Hi-C-constrained physical models of human chromosomes recover functionally-related properties of genome organization

Combining genome-wide structural models with phenomenological data is at the forefront of efforts to understand the organizational principles regulating the human genome. Here, we use chromosome-chromosome contact data as knowledge- based constraints for large-scale three-dimensional models of the human diploid genome. The resulting models remain minimally entangled and acquire several functional features that are observed in vivo and that were never used as input for the model. We find, for instance, that gene-rich, active regions are drawn towards the nuclear center, while gene poor and lamina-associated domains are pushed to the periphery. These and other properties persist upon adding local contact constraints, suggesting their compatibility with non-local constraints for the genome organization. The results show that suitable combinations of data analysis and physical modelling can expose the unexpectedly rich functionally-related properties implicit in chromosome-chromosome contact data. Specific directions are suggested for further developments based on combining experimental data analysis and genomic structural modelling.

biophysics

GW-CALL: Accurate Genome-Wide Variant Caller

The main challenge in reliable variant calling using DNA reads is to extract information from reads mappable to multiple locations on the reference genome. Conventional approaches ignore these reads and rely on reads mappable uniquely to the reference genome. These approaches fail to perform satisfactorily in variant calling within repeat regions which are abundant in many species including homo sapiens. This, in turn, lowers the reliability of any downstream analysis including poor performance in genome-wide association studies. GW-CALL, a fast and accurate variant caller, is proposed. GW-CALL exploits information of all reads in a genome-wide decision making process. In particular, it partitions the genome into several independent regions called clusters and incorporates an efficient algorithm to use all reads belonging to a cluster in calling variants within that cluster.\n\nAvailabilityGW-CALL is implemented in C++ and is freely available at URL: brl.ce.sharif.edu/gwcall.

Bioinformatics

Semi-automated genome annotation using epigenomic data and Segway

Biochemical techniques measure many individual properties of chromatin along the genome. These properties include DNA accessibility (measured by DNase-seq) and the presence of individual transcription factors and histone modifications (measured by ChIP-seq). Segway is software that transforms multiple datasets on chromatin properties into a single annotation of the genome that a biologist can more easily interpret. This protocol describes how to use Segway to annotate the genome, starting with reads from a ChIP-seq experiment. It includes pre-processing of data, training the Segway model, annotating the genome, assigning biological meanings to labels, and visualizing the annotation in a genome browser.

bioinformatics

Ribbon: Visualizing complex genome alignments and structural variation

To the EditorVisualization has played an extremely important role in the current genomic revolution to inspect and understand variants, expression patterns, evolutionary changes, and a number of other relationships1-3. However, most of the information in read-to-reference or genome-genome alignments is lost for structural variations in the one-dimensional views of most genome browsers showing only reference coordinates. Instead, structural variations captured by long reads or assembled contigs often need more context to understand, including alignments and other genomic information from multiple chromosomes.

bioinformatics