Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

From raw reads to trees: Whole genome SNP phylogenetics across the tree of life

Next-generation sequencing is increasingly being used to examine closely related organisms. However, while genome-wide single nucleotide polymorphisms (SNPs) provide an excellent resource for phylogenetic reconstruction, to date evolutionary analyses have been performed using different ad hoc methods that are not often widely applicable across different projects. To facilitate the construction of robust phylogenies, we have developed a method for genome-wide identification/characterization of SNPs from sequencing reads and genome assemblies. Our phylogenetic and molecular evolutionary (PhaME) analysis software is unique in its ability to take reads and draft/complete genome(s) as input, derive core genome alignments, identify SNPs, construct phylogenies and perform evolutionary analyses. Several examples using genomes and read datasets for bacterial, eukaryotic and viral linages demonstrate the broad and robust functionality of PhaME. Furthermore, the ability to incorporate raw metagenomic reads from clinical samples with suspected infectious agents shows promise for the rapid phylogenetic characterization of pathogens within complex samples.

Bioinformatics

Genomic analysis reveals major determinants of cis-regulatory variation in Capsella grandiflora

Understanding the causes of cis-regulatory variation is a long-standing aim in evolutionary biology. Although cis-regulatory variation has long been considered important for adaptation, we still have a limited understanding of the selective importance and genomic determinants of standing cis-regulatory variation. To address these questions, we studied the prevalence, genomic determinants and selective forces shaping cis-regulatory variation in the outcrossing plant Capsella grandiflora. We first identified a set of 1,010 genes with common cis-regulatory variation using analyses of allele-specific expression (ASE). Population genomic analyses of whole-genome sequences from 32 individuals showed that genes with common cis-regulatory variation are 1) under weaker purifying selection and 2) undergo less frequent positive selection than other genes. We further identified genomic determinants of cis-regulatory variation. Gene-body methylation (gbM) was a major factor constraining cis-regulatory variation, whereas presence of nearby TEs and tissue specificity of expression increased the odds of ASE. Our results suggest that most common cis-regulatory variation in C. grandiflora is under weak purifying selection, and that gene-specific functional constraints are more important for the maintenance of cis-regulatory variation than genome-scale variation in the intensity of selection. Our results agree with previous findings that suggest TE silencing affects nearby gene expression, and provide novel evidence for a link between gbM and cis-regulatory constraint, possibly reflecting greater dosage-sensitivity of body-methylated genes. Given the extensive conservation of gene-body methylation in flowering plants, this suggests that gene-body methylation could be an important predictor of cis-regulatory variation in a wide range of plant species.

Evolutionary Biology

chromPlot: visualization of genomic data in chromosomal context

Summary: Visualizing genomic data in chromosomal context can help detecting errors in data generation or analysis and can suggest new hypotheses to be tested. Here we report a new tool for displaying large and diverse genomic data in idiograms of one or multiple chromosomes. The package is implemented in R so that visualization can be easily integrated with its numerous packages for processing genomic data. It supports simultaneous visualization of multiples tracks of data, each of potentially different nature. Large genomic regions such as QTLs or synteny tracts may be shown along histograms of number of genes, genetic variants, or any other type of genomic element. Tracks can also contain values for continuous or categorical variables and the user can choose among points, points connected by lines, line segments, barplots or histograms for representing data. chromPlot reads data from tables in BED format which are imported in R using its builtin functions. The information necessary to draw chromosomes for mouse and human is included with the package. Chromosomes for other organisms are downloaded automatically from the Ensembl website or can be provided by the user. We present common use cases here, and a full tutorial is included as the packages's vignette.\n\nAvailability: chromPlot is distributed under a GLP2 licence at Genomed Lab: http://genomed.med.uchile.cl.\n\nContact: raverdugo@u.uchile.cl

Bioinformatics

A Model of Avian Genome Evolution

A model of genome evolution is proposed. Based on several general assumptions the evolutionary theory of a genome is formulated. Both the deterministic classical equation and the stochastic quantum equation are proposed. The classical equation is written in a form of of second-order differential equations on nucleotide frequencies varying in time. It is proved that the evolutionary equation can be put in a form of the least action principle and the latter can be used for obtaining the quantum generalization of the evolutionary law. The wave equation and uncertainty relation for the quantum evolution are deduced logically. Two fundamental constants of time dimension, the quantization constant and the evolutionary inertia, are introduced for characterizing the genome evolution. During speciation the large-scale rapid change of nucleotide frequency makes the evolutionary inertia of the dynamical variables of the genome largely decreasing or losing. That leads to the occurrence of quantum phase of the evolution. The observed smooth/sudden evolution is interpreted by the alternating occurrence of the classical and quantum phases. In this theory the probability of new-species formation is calculable from the first-principle. To deep the discussions we consider avian genome evolution as an example. More concrete forms on the assumed potential in fundamental equations, namely the diversity and the environmental potential, are introduced. Through the numerical calculations we found that the existing experimental data on avian macroevolution are consistent with our theory. Particularly, the law of the rapid post-Cretaceous radiation of neoavian birds can be understood in the quantum theory. Finally, the present work shows the quantum law may be more general than thought, since it plays key roles not only in atomic physics, but also in genome evolution.

Evolutionary Biology

Tagmentation-Based Mapping (TagMap) of Mobile DNA Genomic Insertion Sites

Multiple methods have been introduced over the past 30 years to identify the genomic insertion sites of transposable elements and other DNA elements that integrate into genomes. However, each of these methods suffer from limitations that can frustrate attempts to map multiple insertions in a single genome and to map insertions in genomes of high complexity that contain extensive repetitive DNA. I introduce a new method for transposon mapping that is simple to perform, can accurately map multiple insertions per genome, and generates long sequence \"reads\" that facilitate mapping to complex genomes. The method, called TagMap, for Tagmentation-based Mapping, relies on a modified Tn5 tagmentation protocol with a single tagmentation adaptor followed by PCR using primers specific to the tranposable element and the adaptor sequence. Several minor modifications to normal tagmentation reagents and protocols allow easy and rapid preparation of TagMap libraries. Short read sequencing starting from the adaptor sequence generates oriented reads that flank and are oriented toward the transposable element insertion site. The convergent orientation of adjacent reads at the insertion site allows straightforward prediction of the precise insertion site(s). A Linux shell script is provided to identify insertion sites from fastq files.

Molecular Biology

Natural selection driven by DNA binding proteins shapes genome-wide motif statistics

Ectopic DNA binding by transcription factors and other DNA binding proteins can be detrimental to cellular functions and ultimately to organismal fitness. The frequency of protein-DNA binding at non-functional sites depends on the global composition of a genome with respect to all possible short motifs, or k-mer words. To determine whether weak yet ubiquitous protein-DNA interactions could exert significant evolutionary pressures on genomes, we correlate in vitro measurements of binding strengths on all 8-mer words from a large collection of transcription factors, in several different species, against their relative genomic frequencies. Our analysis reveals a clear signal of purifying selection to reduce the large number of weak binding sites genome-wide. This evolutionary process, which we call global selection, has a detectable hallmark in that similar words experience similar evolutionary pressure, a consequence of the biophysics of protein-DNA binding. By analyzing a large collection of genomes, we show that global selection exists in all domains of life, and operates through tiny selective steps, maintaining genomic binding landscapes over long evolutionary timescales.

Evolutionary Biology

Peculiar hybrid genomes of devastating plant pests promote plasticity in the absence of sex and meiosis

Root-knot nematodes (genus Meloidogyne) show an intriguing diversity of reproductive modes ranging from obligatory sexual to fully asexual reproduction. Intriguingly, the most damaging species to the world agriculture are those that reproduce without meiosis and without sex. To understand this parasitic success despite the absence of sex and genetic exchanges, we have sequenced and assembled the genomes of 3 obligatory ameiotic asexual Meloidogyne species and have compared them to those of meiotic relatives with facultative or obligatory asexual reproduction. Our comparative genomic analysis shows that obligatory asexual root-knot nematodes have a higher abundance of transposable elements (TE) compared to the facultative sexual and contain duplicated regions with a high within-species average nucleotide divergence of 8%. Phylogenomic analysis of the genes present in these duplicated regions suggests that they originated from multiple hybridization events. The average nucleotide divergence in the coding portions between duplicated regions is ~5-6 % and we detected diversifying selection between the corresponding gene copies. Genes under diversifying selection covered a wide spectrum of predicted functional categories which suggests a high impact of the genome structure at the functional level. Contrasting with high within-species nuclear genome divergence, mitochondrial genome divergence between the three ameiotic asexuals was very low, suggesting that these putative hybrids share a recent common maternal donor lineage. The intriguing parasitic success of mitotic root-knot nematodes in the absence of sex may be partly explained by TE-rich composite genomes resulting from multiple allo-polyploidization events and promoting plasticity in the absence of sex.

Evolutionary Biology

RefSoil: A reference database of soil microbial genomes

A database of curated genomes is needed to better assess soil microbial communities and their processes associated with differing land management and environmental impacts. Interpreting soil metagenomic datasets with existing sequence databases is challenging because these datasets are biased towards medical and biotechnology research and can result in misleading annotations. We have curated a database of 922 genomes of soil-associated organisms (888 bacteria and 34 archaea). Using this database, we evaluated phyla and functions that are enriched in soils as well as those that may be underrepresented in RefSoil. Our comparison of RefSoil to soil amplicon datasets allowed us to identify targets that if cultured or sequenced would significantly increase the biodiversity represented within RefSoil. To demonstrate the opportunities to access these underrepresented targets, we employed single cell genomics in a pilot experiment to sequence 14 genomes. This effort demonstrates the value of RefSoil in the guidance of future research efforts and the capability of single cell genomics as a practical means to fill the existing genomic data gaps.

Ecology

Genome reduction in an abundant and ubiquitous soil bacterial lineage

Although bacteria within the Verrucomicrobia phylum are pervasive in soils around the world, they are underrepresented in both isolate collections and genomic databases. Here we describe a single verrucomicrobial phylotype within the class Spartobacteria that is not closely related to any previously described taxa. We examined >1000 soils and found this spartobacterial phylotype to be ubiquitous and consistently one of the most abundant soil bacterial phylotypes, particularly in grasslands, where it was typically the most abundant phylotype. We reconstructed a nearly complete genome of this phylotype from a soil metagenome for which we propose the provisional name Candidatus Udaeobacter copiosus. The Ca. U. copiosus genome is unusually small for soil bacteria, estimated to be only 2.81 Mbp compared to the predicted effective mean genome size of 4.74 Mbp for soil bacteria. Metabolic reconstruction suggests that Ca. U. copiosus is an aerobic chemoorganoheterotroph with numerous amino acid and vitamin auxotrophies. The large population size, relatively small genome and multiple putative auxotrophies characteristic of Ca. U. copiosus suggests that it may be undergoing streamlining selection to minimize cellular architecture, a phenomenon previously thought to be restricted to aquatic bacteria. Although many soil bacteria need relatively large, complex genomes to be successful in soil, Ca. U. copiosus appears to have identified an alternate strategy, sacrificing metabolic versatility for efficiency to become dominant in the soil environment.

Microbiology

Lepbase: the Lepidopteran genome database

As the generation and use of genomic datasets is becoming increasingly common in all areas of biology, the need for resources to collate, analyse and present data from independent (Tier 1) species-level genome projects into well supported clade-oriented (Tier 2) databases and provide a mechanism for these data to be propagated to pan-taxonomic (Tier 3) databases is becoming more pressing. Lepbase is a Tier 2 genomic resource for the Lepidoptera, supporting a research community using genomic approaches to understand evolution, speciation, olfaction, behaviour and pesticide resistance in a wide range of target species. Lepbase offers a core set of tools to make genomic data widely accessible including an Ensembl genome browser, text and sequence homology searches and bulk downloads of consistently presented and formatted datasets. As a part of the taxonomic community that we serve, we are working directly with Lepidoptera researchers to prioritise analyses and add tools that will be of most value to current research questions.

Bioinformatics

A Novel Family of Genomics Islands Across Multiple Species of Streptococcus

The genus Streptococcus is one of the most genomically diverse and important human and agricultural pathogens. The acquisition of genomic islands (GIs) plays a central role in adaptation to new hosts in the genus pathogens. The research presented here employs a comparative genomics approach to define a novel family of GIs in the genus Streptococcus which also appears across strains of the same species. Specifically, we identified 9 Streptococcus genomes out of 67 sequenced genomes analyzed, and we termed these as 15bp Streptococcus genomic islands, or 15SGIs, including i) insertion adjacent to the 3 end of ribosome l7/l12 gene, ii) large inserts of horizontally acquired DNA, and iii) the presence of mobility genes (integrase) and replication initiators. We have identified a novel family of 15SGIs and seems to be important in species differentiation and adaptation to new hosts. It plays an important role during strain evolution in the genus Streptococcus.

Microbiology

CASTOR: A machine learning platform for reproducible viral genome classification

MotivationAdvances in cloning and sequencing technology yielded a massive number of genome of virus strains. The classification and annotation of these genomes constitute important assets in the discovery of genomic variability, taxonomic characteristics and disease mechanisms. Existing classification methods are often designed for a well-studied virus. Thus, the viral comparative genomic studies could benefit from more generic, fast and accurate tools for classifying and typing newly sequenced strains of diverse virus families.\n\nResultsHere, we introduce a fast, accurate and generic virus classification platform, CASTOR, based on a machine learning approach. CASTOR is inspired by a well-known technique in molecular biology: Restriction Fragment Length Polymorphism (RFLP). It simulates the restriction digestion of genomic material by different enzymes into fragments in-silico. It uses two metrics to construct feature vectors for machine learning algorithms in the classification step. We benchmark CASTOR for the classification of distinct datasets of Human Papillomaviruses (HPV), Hepatitis B Viruses (HBV) and Human Immunodeficiency viruses (HIV). Results reveal true positive rates of 99%, 99% and 98% for HPV Alpha species, HBV genotyping and HIV M group subtyping respectively. Furthermore, CASTOR shows a competitive performance compare to well-known HIV-specific classifier REGA and COMET on whole genome and pol fragments. With such prediction rates, genericity and robustness, as well as rapidity, such approach could constitute a reference in large-scale virus studies. Finally, we developed the CASTOR web platform for open access and reproducible viral machine learning classifiers.\n\nAvailabilityhttp://castor.bioinfo.uqam.ca\n\nContactdiallo.abdoulaye@uqam.ca

bioinformatics

Ensembl Core Software Resources: storage and programmatic access for DNA sequence and genome annotation

The Ensembl software resources are a stable infrastructure to store, access and manipulate genome assemblies and their functional annotations. The Ensembl \"Core\" database and Application Programming Interface (API) was our first major piece of software infrastructure and remains at the centre of all of our genome resources. Since its initial design more than fifteen years ago, the number of publicly available genomic, transcriptomic and proteomic datasets has grown enormously, accelerated by continuous advances in DNA sequencing technology. Initially intended to provide annotation for the reference human genome, we have extended our framework to support the genomes of all species as well as richer assembly models. Cross-referenced links to other informatics resources facilitate searching our database with a variety of popular identifiers such as UniProt and RefSeq. Our comprehensive and robust framework storing a large diversity of genome annotations in one location serves as a platform for other groups to generate and maintain their own tailored annotation. Our databases and APIs are publicly available and all of our source code is released with a permissive Apache v2.0 licence at http://github.com/Ensembl.

bioinformatics

Scalable parameter estimation for genome-scale biochemical reaction networks

Mechanistic mathematical modeling of biochemical reaction networks using ordinary differential equation (ODE) models has improved our understanding of small-and medium-scale biological processes. While the same should in principle hold for large-and genome-scale processes, the computational methods for the analysis of ODE models which describe hundreds or thousands of biochemical species and reactions are missing so far. While individual simulations are feasible, the inference of the model parameters from experimental data is computationally too intensive. In this manuscript, we evaluate adjoint sensitivity analysis for parameter estimation in large scale biochemical reaction networks. We present the approach for time-discrete measurement and compare it to state-of-the-art methods used in systems and computational biology. Our comparison reveals a significantly improved computational efficiency and a superior scalability of adjoint sensitivity analysis. The computational complexity is effectively independent of the number of parameters, enabling the analysis of large-and genome-scale models. Our study of a comprehensive kinetic model of ErbB signaling shows that parameter estimation using adjoint sensitivity analysis requires a fraction of the computation time of established methods. The proposed method will facilitate mechanistic modeling of genome-scale cellular processes, as required in the age of omics.\n\nAuthor SummaryIn this manuscript, we introduce a scalable method for parameter estimation for genome-scale biochemical reaction networks. Mechanistic models for genome-scale biochemical reaction networks describe the behavior of thousands of chemical species using thousands of parameters. Standard methods for parameter estimation are usually computationally intractable at these scales. Adjoint sensitivity based approaches have been suggested to have superior scalability but any rigorous evaluation is lacking. We implement a toolbox for adjoint sensitivity analysis for biochemical reaction network which also supports the import of SBML models. We show by means of a set of benchmark models that adjoint sensitivity based approaches unequivocally outperform standard approaches for large-scale models and that the achieved speedup increases with respect to both the number of parameters and the number of chemical species in the model. This demonstrates the applicability of adjoint sensitivity based approaches to parameter estimation for genome-scale mechanistic model. The MATLAB toolbox implementing the developed methods is available from http://ICB-DCM.github.io/AMICI/.

systems biology

A Complete Logical Approach to Resolve the Evolution and Dynamics of Mitochondrial Genome in Bilaterians

A new method of genomic maps analysis based on formal logic is described. The purpose of the method is to 1) use mitochondrial genomic organisation of current taxa as datasets 2) calculate mutational steps between all mitochondrial gene arrangements and 3) reconstruct phylogenetic relationships according to these calculated mutational steps within a dendrogram under the assumption of maximum parsimony. Unlike existing methods mainly based on the probabilistic approach, the main strength of this new approach is that it calculates all the exact tree solutions with completeness and provides logical consequences as very robust results. Moreover, the method infers all possible hypothetical ancestors and reconstructs character states for all internal nodes (ancestors) of the trees. We started by testing the method using the deuterostomes as a study case. Then, with sponges as an outgroup, we investigated the mutational network of mitochondrial genomes of 47 bilaterian phyla and emphasised the peculiar case of chaetognaths. This pilot work showed that the use of formal logic in a hypothetico-deductive background such as phylogeny (where experimental testing of hypotheses is impossible) is very promising to explore mitochondrial gene rearrangements in deuterostomes and should be applied to many other bilaterian clades.\n\nAuthor SummaryInvestigating how recombination might modify gene arrangements during the evolution of metazoans has become a routine part of mitochondrial genome analysis. In this paper, we present a new approach based on formal logic that provides optimal solutions in the genome rearrangement field. In particular, we improve the sorting by including all rearrangement events, e.g., transposition, inversion and reverse transposition. The problem we face with is to find the most parsimonious tree(s) explaining all the rearrangement events from a common ancestor to all the descendants of a given clade (hereinafter PHYLO problem). So far, a complete approach to find all the correct solutions of PHYLO is not available. Formal logic provides an elegant way to represent and solve such an NP-hard problem. It has the benefit of correctness, completeness and allows the understanding of the logical consequences (results true for all solutions found). First, one must define PHYLO (axiomatisation) with a set of logic formulas or constraints. Second, a model generator calculates all the models, each model being a solution of PHYLO. Several complete model generators are available but a recurring difficulty is the computation time when the data set increases. When the search of a solution takes exponential time, two computing strategies are conceivable: an incomplete but fast algorithm that does not provide the optimal solution (for example, use local improvements from an initial random solution) or a complete - and thus not efficient - algorithm on a smaller tractable dataset. While the large amount of genes found in the nuclear genome strongly limits our possibility to use of formal logic with any conventional computer, we show in our paper that, for bilaterian mtDNAs, all the correct solutions can be found in a reasonable time due to the small number of genes.

bioinformatics

Genome Graphs

There is increasing recognition that a single, monoploid reference genome is a poor universal reference structure for human genetics, because it represents only a tiny fraction of human variation. Adding this missing variation results in a structure that can be described as a mathematical graph: a genome graph. We demonstrate that, in comparison to the existing reference genome (GRCh38), genome graphs can substantially improve the fractions of reads that map uniquely and perfectly. Furthermore, we show that this fundamental simplification of read mapping transforms the variant calling problem from one in which many non-reference variants must be discovered de-novo to one in which the vast majority of variants are simply re-identified within the graph. Using standard benchmarks as well as a novel reference-free evaluation, we show that a simplistic variant calling procedure on a genome graph can already call variants at least as well as, and in many cases better than, a state-of-the-art method on the linear human reference genome. We anticipate that graph-based references will supplant linear references in humans and in other applications where cohorts of sequenced individuals are available.

bioinformatics

Tracking multiple genomic elements using correlative CRISPR imaging and sequential DNA FISH

Live imaging of genome has offered important insights into the dynamics of the genome organization and gene expression. The demand to image simultaneously multiple genomic loci has prompted a flurry of exciting advances in multi-color CRISPR imaging, although color-based multiplexing is limited by the need for spectrally distinct fluorophores. Here we introduce an approach to achieve highly multiplexed live recording via correlative CRISPR imaging and sequential DNA fluorescence in situ hybridization (FISH). This approach first performs one-color live imaging of multiple genomic loci and then uses sequential rounds of DNA FISH to determine the loci identity. We have optimized the FISH protocol so that each round is complete in 1 min, demonstrating the identification of 7 genomic elements and the capability to sustain reversible staining and washing for up to 20 rounds. We have also developed a correlation-based algorithm to faithfully register live and FISH images. Our approach keeps the rest of the color palette open to image other cellular phenomena of interest, as demonstrated by our simultaneous live imaging of genomic loci together with a cell cycle reporter. Furthermore, the algorithm to register faithfully between live and fixed imaging is directly transferrable to other systems such as multiplex RNA imaging with RNA-FISH and multiplex protein imaging with antibody-staining.

biophysics

Genome-wide signatures of genetic variation within and between populations - a comparative perspective

Genome-wide screens of genetic variation can reveal signatures of population-specific selection implicated in adaptation and speciation. Yet, unrelated processes such as linked selection arising as a consequence of genome architecture can generate comparable signatures across taxa. To investigate prevalence and phylogenetic stability of linked selection, we took a comparative approach utilizing population-level data from 444 re-sequenced genomes of three avian clades spanning 50 million years of evolution. Levels of nucleotide diversity ({pi}),population-scaled recombination rate ({rho}), genetic differentiation (FST, PBS) and sequence divergence (Dxy) were remarkably similar in syntenic genomic regions across clades. Elevated local genetic differentiation was associated with inferred centromere and sub-telomeric regions. Our results support a role of linked selection shaping genome-wide heterogeneity in genetic diversity within and between clades. The long-term conservation of diversity landscapes and stable association with genomic features make the outcome of this evolutionary process in part predictable.

evolutionary biology