Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Metabolic adaptations underlying genome flexibility in prokaryotes

Even across genomes of the same species, prokaryotes exhibit remarkable flexibility in gene content. We do not know whether this flexible or \"accessory\" content is mostly neutral or adaptive, largely due to the lack of explicit analyses of accessory gene function. Here, across 96 diverse prokaryotic species, I show that a considerable fraction (~40%) of accessory genomes harbours beneficial metabolic functions. These functions take two forms: (1) they significantly expand the biosynthetic potential of individual strains, and (2) they help reduce strain-specific metabolic auxotrophies via intra-species metabolic exchanges. I find that the potential of both these functions increases with increasing genome flexibility. Together, these results are consistent with a significant adaptive role for prokaryotic pangenomes.\n\nAuthor SummaryRecent and rapid advancements in genome sequencing technologies have revealed key insights into the world of bacteria and archaea. One puzzling aspect uncovered by these studies is the following: genomes of the same species can often look very different. Specifically, some \"core\" genes are maintained across all intraspecies genomes, but many \"accessory\" genes differ between strains. A major ongoing debate thus asks: do most of these accessory genes provide a benefit to different strains, and if so, in what form? In this study, I suggest that the answer is \"yes, through metabolic interactions\". I show that many accessory genes provide significant metabolic advantages to different strains in different conditions. I achieve this by explicitly conducting a large-scale systematic analysis of 1,339 genomes across 96 diverse species of bacteria and archaea. A surprising prediction of this study that in many ecological niches, co-occurring strains of the same species may help each other survive by exchanging metabolites exclusively produced by these different accessory genes. More pronounced gene differences lead to more underlying metabolic advantages.

ecology

Genome-wide disruption of DNA methylation by 5-aza-2’-deoxycytidine in the parasitioid wasp Nasonia vitripennis

DNA methylation of cytosine residues across the genome influences how genes and phenotypes are regulated in a wide range of organisms. As such, understanding the role of DNA methylation and other epigenetic mechanisms has become very much a part of mapping genotype to phenotype, a major question in evolutionary biology. Ideally, we would like to manipulate DNA methylation patterns on a genome-wide scale, to help us to elucidate the role that epigenetic modifications play in phenotypic expression. Recently, the demethylating agent 5-aza-2-deoxycytidine (5-aza-dC; commonly used in the epigenetic treatment of certain cancers), has been deployed to explore the epigenetic regulation of a number of traits of interest to evolutionary ecologists, including facultative sex allocation in the parasitoid wasp Nasonia vitripennis. In a recent study, we showed that treatment with 5-aza-dC did not ablate the facultative sex allocation response in Nasonia, but shifted the patterns of sex allocation in a way predicted by genomic conflict theory. This was the first (albeit indirect) experimental evidence for genomic conflict over sex allocation facilitated by DNA methylation. However, that work lacked direct evidence of the effects of 5-aza-dC on DNA methylation, and indeed the effect of the chemical has since been questioned in Nasonia. Here, using whole-genome bisulphite sequencing of more than 4 million CpGs, across more than 11,000 genes, we demonstrate unequivocally that 5-aza-dC disrupts methylation on a large scale across the Nasonia vitripennis genome. We show that the disruption can lead to both hypo- and hyper-methylation, may vary across tissues and time of sampling, and that the effects of 5-aza-dC are context- and sequence specific. We conclude that 5-aza-dC does indeed have the potential to be repurposed as a tool for studying the role of DNA methylation in evolutionary ecology, whilst many details of its action remain to be discovered. Author SummaryShedding light on the mechanistic basis of phenotypes is a major aim in the field of evolutionary biology. If we understand how phenotypes are controlled at the molecular level, we can begin to understand how evolution has shaped that phenotype and conversely, how genetic architecture may constrain trait evolution. Epigenetic markers (such as DNA methylation) also influence phenotypic expression by regulating how and when genes are expressed. Recently, 5-aza-2-deoxycytidine (5-aza-dC), a hypomethylating agent used in the treatment of certain cancers, has been used to explore the epigenetic regulation of traits of interest to evolutionary ecologists. Previously, we used 5-aza-dC to validate a role for DNA methylation in facultative sex allocation behaviour in the parasitoid wasp Nasonia vitripennis. However, the direct effects of the chemical were not examined at that point and its efficacy in insects was questioned. Here, we demonstrate that 5-aza-dC disrupts DNA methylation on a genome-wide scale in a context- and sequence-specific manner and results in both hypo- and hyper-methylation. Our work demonstrates that 5-aza-dC has the potential to be repurposed as a tool for studying the role of DNA methylation in phenotypic expression.

evolutionary biology

Genomic signatures of extensive inbreeding in Isle Royale wolves, a population on the threshold of extinction

The observation that small, isolated populations often suffer reduced fitness as a result of inbreeding depression has guided conservation theory and practice for decades. However, investigating the genome-wide dynamics associated with inbreeding depression in natural populations is only now feasible with relatively inexpensive sequencing technology and annotated reference genomes. To characterize the genome-wide effects of intense inbreeding and isolation, we sequenced complete genomes from an iconic inbred population, the gray wolves (Canis lupus) of Isle Royale. Through comparison with other wolf genomes from a variety of demographic histories, we found that Isle Royale wolf genomes contain extensive runs of homozygosity, but neither the overall level of heterozygosity nor the number of deleterious variants per genome were reliable predictors of inbreeding depression. These findings are consistent with the hypothesis that severe inbreeding depression results from increased homozygosity of strongly deleterious recessive mutations, which are more prevalent in historically large source populations. Our results have particular relevance in light of the recently proposed reintroduction of wolves to Isle Royale, as well as broader implications for management of genetic variation in the fragmented landscape of the modern world.

evolutionary biology

Selective advantages favour high genomic AT-contents in intracellular elements

Extrachromosomal genetic elements generally exhibit increased AT-contents relative to their hosts DNA. The AT-bias of endosymbiotic genomes is commonly explained by neutral evolutionary processes. Here we show experimentally that an increased AT-content of host-dependent elements can be selectively favoured on the host level. Manipulating the nucleotide composition of bacterial cells by introducing A+T-or G+C-rich plasmids, we demonstrate that cells containing GC-rich plasmids are less fit than cells containing AT-rich plasmids. Moreover, the cost of GC-rich elements could be compensated by providing G+C-, but not A+T-precursors, thus linking the observed fitness effects to the cytoplasmic availability of nucleotides. Our work identifies selection as a strong evolutionary force that drives the genomes of intracellular genetic elements toward higher A+T contents.\n\nAuthor SummaryGenomes of endosymbiotic bacteria are commonly more AT-rich than the ones of their free-living relatives. Interestingly, genomes of other intracellular elements like plasmids or bacteriophages also tend to be richer in AT than the genomes of their hosts. The AT-bias of endosymbiotic genomes is commonly explained by neutral evolutionary processes. However, since A+T nucleotides are both more abundant and energetically less expensive than G+C nucleotides, an alternative explanation is that selective advantages drive the nucleotide composition of intracellular elements. Here we provide strong experimental evidence that intracellular elements, whose genome is more AT-rich than the genome of the host, are selectively favored on the host level. Thus, our results emphasize the importance of selection for shaping the DNA base composition of extrachromosomal genetic elements.

evolutionary biology

The Clinical Imperative for Inclusivity: Race, Ethnicity, and Ancestry (REA) in Genomics

The Clinical Genome Resource (ClinGen) Ancestry and Diversity Working Group highlights the need to develop guidance on race, ethnicity, and ancestry (REA) data collection and use in clinical genomics. We present quantitative and qualitative evidence to characterize: 1) acquisition of REA data via clinical laboratory requisition forms, and 2) information disparity across populations in the Genome Aggregation Database (gnomAD) at clinically relevant sites as determined by variants in ClinVar. Our requisition form analysis showed substantial heterogeneity in clinical laboratory ascertainment of REA, as well as marked incongruity among terms used to define REA categories. There was also striking disparity across REA populations in the amount of information available about variants at clinically relevant sites in gnomAD. European ancestral populations constituted the majority of observations (55.8%), allele counts (59.7%), and private alleles (56.1%) in gnomAD at 550 loci with \"pathogenic\" and \"likely pathogenic\" expert-reviewed variants in ClinVar. Our findings highlight the importance of implementing and supporting programs to increase diversity in genome sequencing and clinical genomics, as well as measuring uncertainty around population-level datasets that are used in variant interpretation. Finally, we suggest the need for a standardized REA data collection framework to be developed and adopted across clinical genomics.

genomics

Bayesian inference of infectious disease transmission from whole genome sequence data

Genomics is increasingly being used to investigate disease outbreaks, but an important question remains unanswered - how well do genomic data capture known transmission events, particularly for pathogens with long carriage periods or large within-host population sizes? Here we present a novel Bayesian approach to reconstruct densely-sampled outbreaks from genomic data whilst considering within-host diversity. We infer a time-labelled phylogeny using BEAST, then infer a transmission network via a Monte-Carlo Markov Chain. We find that under a realistic model of within-host evolution, reconstructions of simulated outbreaks contain substantial uncertainty even when genomic data reflect a high substitution rate. Reconstruction of a real-world tuberculosis outbreak displayed similar uncertainty, although the correct source case and several clusters of epidemiologically linked cases were identified. We conclude that genomics cannot wholly replace traditional epidemiology, but that Bayesian reconstructions derived from sequence data may form a useful starting point for a genomic epidemiology investigation.

Genomics

Probabilities of Fitness Consequences for Point Mutations Across the Human Genome

We describe a novel computational method for estimating the probability that a point mutation at each position in a genome will influence fitness. These fitness consequence (fit-Cons) scores serve as evolution-based measures of potential genomic function. Our approach is to cluster genomic positions into groups exhibiting distinct \"fingerprints\" based on high-throughput functional genomic data, then to estimate a probability of fitness consequences for each group from associated patterns of genetic polymorphism and divergence. We have generated fitCons scores for three human cell types based on public data from EN-CODE. Compared with conventional conservation scores, fitCons scores show considerably improved prediction power for cis-regulatory elements. In addition, fitCons scores indicate that 4.2-7.5% of nucleotides in the human genome have influenced fitness since the human-chimpanzee divergence, and, in contrast to several recent studies, they suggest that recent evolutionary turnover has had limited impact on the functional content of the genome.

Genomics

The Sea Lamprey Meiotic Map Resolves Ancient Vertebrate Genome Duplications

Gene and genome duplications serve as an important reservoir of material for the evolution of new biological functions. It is generally accepted that many genes present in vertebrate genomes owe their origin to two whole genome duplications that occurred deep in the ancestry of the vertebrate lineage. However, details regarding the timing and outcome of these duplications are not well resolved. We present high-density meiotic and comparative genomic maps for the sea lamprey, a representative of an ancient lineage that diverged from all other vertebrates approximately 550 million years ago. Linkage analyses yielded a total of 95 linkage groups, similar to the estimated number of germline chromosomes (1N [~] 99), spanning a total of 5,570.25 cM. Comparative mapping data yield strong support for one ancient whole genome duplication but do not strongly support a hypothetical second event. Rather, these comparative maps reveal several evolutionary independent segmental duplications occurring over the last 600+ million years of chordate evolution. This refined history of vertebrate genome duplication should permit more precise investigations into the evolution of vertebrate gene functions.

Genomics

Estimating gene expression and codon specific translational efficiencies, mutation biases, and selection coefficients from genomic data alone.

Extracting biologically meaningful information from the continuing flood of genomic data is a major challenge in the life sciences. Codon usage bias (CUB) is a general feature of most genomes and is thought to reflect the effects of both natural selection for efficient translation and mutation bias. Here we present a mechanistically interpretable, Bayesian model (ROC SEMPPR) to extract biologically meaningful information from patterns of CUB within a genome. ROC SEMPPR, is grounded in population genetics and allows us to separate the contributions of mutational biases and natural selection against translational inefficiency on a gene by gene and codon by codon basis. Until now, the primary disadvantage of similar approaches was the need for genome scale measurements of gene expression. Here we demonstrate that it is possible to both extract accurate estimates of codon specific mutation biases and translational efficiencies while simultaneously generating accurate estimates of gene expression, rather than requiring such information. We demonstrate the utility of ROC SEMPPR using the S. cerevisiae S288c genome. When we compare our model fits with previous approaches we observe an exceptionally high agreement between estimates of both codon specific parameters and gene expression levels ({rho} > 0.99 in all cases). We also observe strong agreement between our parameter estimates and those derived from alternative datasets. For example, our estimates of mutation bias and those from mutational accumulation exper-iments are highly correlated ({rho} = 0.95). Our estimates of codon specific translational inefficiencies are tRNA copy number based estimates of ribosome pausing time ({rho} = 0.64), and mRNA and ribosome profiling footprint based estimates of gene expression ({rho} = 0.53 - 0.74) are also highly correlated, thus supporting the hypothesis that selection against translational inefficiency is an important force driving the evolution of CUB. Surprisingly, we find that for particular amino acids, codon usage in highly expressed genes can still be largely driven by mutation bias and that failing to take mutation bias into account can lead to the misidentification of an amino acids optimal codon. In conclusion, our method demonstrates that an enormous amount of biologically important information is encoded within genome scale patterns of codon usage, accessing this information does not require gene expression measurements, but instead carefully formulated biologically interpretable models.

Genomics

Uncovering Long Palindromic Sequences in Rice (Oryza sativa subsp. indica) Genome

Because the rice genome has been sequenced entirely, search to find specific features at genome-wide scale is of high importance. Palindromic sequences are important DNA motifs involved in the regulation of different cellular processes and are a potential source of genetic instability. In order to search and study the long palindromic regions in the rice genome \"R\" statistical programming language was used. All palindromes, defined as identical inverted repeats with spacer DNA, could be analyzed and sorted according to their frequency, size, GC content, compact index etc. The results showed that the overall palindrome frequency was high in rice genome (nearly 51000 palindromes), with highest and lowest number of palindromes, respectively belongs to chromosome 1 and 12. Palindrome numbers could well explain the rice chromosome expansion (R2>92%). Average GC content of the palindromic sequences is 42.1%, indicating AT-richness and hence, the low-complexity of palindromic sequences. The results also showed different compact indices of palindromes in different chromosomes (43.2 per cM in chromosome 8 and 34.5 per cM in chromosome 3, as highest and lowest, respectively). The possible application of palindrome identification can be the use in the development of a molecular marker system facilitating some genetic studies such as evaluation of genetic variation and gene mapping and also serving as a useful tool in population structure analysis and genome evolution studies. Based on these results it can be concluded that the rice genome is rich in long palindromic sequences that triggered most variation during evolution.\n\nAvailability: The R scripts used to construct the palindrome sequence library are available upon request.

Genomics

The advent of genome-wide association studies for bacteria

Highlights* The advent of the genome-wide association study (GWAS) approach provides a promising framework for dissecting the genetic basis of bacterial or archaeal phenotypes.\n\n* Bacterial genomes tend to be shaped by stronger positive selection, stronger linkage disequilibrium and stronger population stratification than humans, with implications for GWAS power and resolution.\n\n* An example GWAS in Mycobacterium tuberculosis genomes highlights the potentially confounding effects of linkage disequilibrium and population stratification.\n\n* A comparison of the traditional GWAS approach versus a somewhat orthogonal method based upon evolutionary convergence (phyC) shows strengths and weaknesses of both approaches.\n\nAbstractSignificant advances in sequencing technologies and genome-wide association studies (GWAS) have revealed substantial insight into the genetic architecture of human phenotypes. In recent years, the application of this approach in bacteria has begun to reveal the genetic basis of bacterial host preference, antibiotic resistance, and virulence. Here, we consider relevant differences between bacterial and human genome dynamics, apply GWAS to a global sample of Mycobacterium tuberculosis genomes to highlight the impacts of linkage disequilibrium, population stratification, and natural selection, and finally compare the traditional GWAS against phyC, a contrasting method of mapping genotype to phenotype based upon evolutionary convergence. We discuss strengths and weaknesses of both methods, and make suggestions for factors to be considered in future bacterial GWAS.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=121 SRC=\"FIGDIR/small/016873_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (25K):\norg.highwire.dtl.DTLVardef@431609org.highwire.dtl.DTLVardef@5bad1borg.highwire.dtl.DTLVardef@c1ff30org.highwire.dtl.DTLVardef@58e539_HPS_FORMAT_FIGEXP M_FIG C_FIG

Genomics

The UCSC Genome Browser database: 2016 update

For the past 15 years, the UCSC Genome Browser (http://genome.ucsc.edu/) has served the international research community by offering an integrated platform for viewing and analyzing information from a large database of genome assemblies and their associated annotations. The UCSC Genome Browser has been under continuous development since its inception with new data sets and software features added frequently. Some release highlights of this year include new and updated genome browsers for various assemblies, including bonobo and zebrafish; new gene annotation sets; improvements to track and assembly hub support; and a new interactive tool, the "Data Integrator", for intersecting data from multiple tracks. We have greatly expanded the data sets available on the most recent human assembly, hg38/GRCh38, to include updated gene prediction sets from GENCODE, more phenotype- and disease-associated variants from ClinVar and ClinGen, more genomic regulatory data, and a new multiple genome alignment.

Genomics

Patching holes in the Chlamydomonas genome

The Chlamydomonas genome has been sequenced, assembled and annotated to produce a rich resource for genetics and molecular biology in this well-studied model organism. However, the current reference genome contains ~1000 blocks of unknown sequence ( N-islands), which are frequently placed in introns of annotated gene models. We developed a strategy, using careful bioinformatics analysis of short-sequence cDNA and genomic DNA reads, to search for previously unknown exons hidden within such blocks, and determine the sequence and exon/intron boundaries of such exons. These methods are based on assembly and alignment completely independent of prior reference assembly or reference annotation. Our evidence indicates that ~one-quarter of the annotated intronic N-islands actually contain hidden exons. For most of these our algorithm recovers full exonic sequence with associated splice junctions and exon-adjacent intron sequence, that can be joined to the reference genome assembly and annotated transcript models. These new exons represent de novo sequence generally present nowhere in the assembled genome, and the added sequence can be shown in many cases to greatly improve evolutionary conservation of the predicted encoded peptides. At the same time, our results confirm the purely intronic status for a substantial majority of N-islands annotated as intronic in the reference annotated genome, increasing confidence in this valuable resource.

Genomics

Genomic Bayesian Prediction Model for Count Data with Genotype × Environment Interaction

Genomic tools allow the study of the whole genome and are facilitating the study of genotype-environment combinations and their relationship with the phenotype. However, most genomic prediction models developed so far are appropriate for Gaussian phenotypes. For this reason, appropriate genomic prediction models are needed for count data, since the conventional regression models used on count data with a large sample size (n) and a small number of parameters (p) cannot be used for genomic-enabled prediction where the number of parameters (p) is larger than the sample size (n). Here we propose a Bayesian mixed negative binomial (BMNB) genomic regression model for counts that takes into account genotype by environment (G x E) interaction. We also provide all the full conditional distributions to implement a Gibbs sampler. We evaluated the proposed model using a simulated data set and a real wheat data set from the International Maize and Wheat Improvement Center (CIMMYT) and collaborators. Results indicate that our BMNB model is a viable alternative for analyzing count data.

Genomics

Short template switch events explain mutation clusters in the human genome

Resequencing efforts are uncovering the extent of genetic variation in humans and provide data to study the evolutionary processes shaping our genome. One recurring puzzle in both intra- and inter-species studies is the high frequency of complex mutations comprising multiple nearby base substitutions or insertion-deletions. We devised a generalized mutation model of template switching during replication that extends existing models of genome rearrangement, and used this to study the role of template switch events in the origin of such mutation clusters. Applied to the human genome, our model detects thousands of template switch events during the evolution of human and chimp from their common ancestor, and hundreds of events between two independently sequenced human genomes. While many of these are consistent with the template switch mechanism previously proposed for bacteria but not thought significant in higher organisms, our model also identifies new types of mutations that create short inversions, some flanked by paired inverted repeats. The local template switch process can create numerous complex mutation patterns, including hairpin loop structures, and explains multi-nucleotide mutations and compensatory substitutions without invoking positive selection, complicated and speculative mechanisms, or implausible coincidence. Clustered sequence differences are challenging for mapping and variant calling methods, and we show that detection of mutation clusters with current resequencing methodologies is difficult and many erroneous variant annotations exist in human reference data. Template switch events such as those we have uncovered may have been neglected as an explanation for complex mutations because of biases in commonly used analyses. Incorporation of our model into reference-based analysis pipelines and comparisons of de novo-assembled genomes will lead to improved understanding of genome variation and evolution.

Genomics

Nomadic Lifestyle of Lactobacillus plantarum Revealed by Comparative Genomics of 54 strains Isolated from Different Niches

The ability of many bacteria to adapt to diverse environmental conditions is well known. Recent research has linked the process of bacterial adaptation to a niche to changes in the genome content and size, showing that many bacterial genomes reflect the constraints imposed by their habitat. However, some highly versatile bacteria are found in diverse niches that almost share nothing in common. Lactobacillus plantarum is a lactic acid bacterium that is found in a large variety of niches. With the aim of unravelling the link between genome evolution and ecological versatility of L. plantarum, we analysed the genomes of 54 L. plantarum strains isolated from different environments. Phylogenomic analyses coupled with the study of genetic functional divergence and gene-trait matching analysis revealed a mixed distribution of the strains, which was uncoupled from their environmental origin. Our findings demonstrate the high complexity of L. plantarum evolution, revealing the absence of specific genomic signatures marking adaptations of this species towards the diverse habitats it is associated with. This suggests fundamentally similar and parallel trends of genome evolution in L. plantarum, which occur in a manner that is apparently uncoupled from ecological constraint and reflects the nomadic lifestyle of this species.

Genomics

A precision metric for clinical genome sequencing

A requisite precondition for the application of next-generation sequencing to clinical medicine is the ability to confidently call genotype at each coding/splicing position of every gene of interest. Current gold standard technologies, such as Sanger sequencing and microarrays, allow confident identification of the genomic origin of the DNA of interest. A commonly used minimum standard for the adoption of new technology in medicine is non-inferiority. We developed a metric to quantify the extent to which current sequencing technologies reach this clinical grade reporting standard. This metric, the rationale for which we present here, is defined as the absolute number of base pairs per gene not callable with confidence, as specified by the presence of 20 high quality (Q20) bases from uniquely mapped (mapq>0) reads per locus. To illustrate the utility of this metric, we apply it across data from several commercially available clinical sequencing products. We present specific examples of coverage for genes known to be important for clinical medicine. We derive data from a variety of platforms including whole genome sequencing (Illumina Hiseq and X chemistry) and exome capture (including medically optimized capture from Agilent, Baylor Clinical Lab, and Personalis). We observe that compared to whole genomes (with {small tilde}30x average coverage), augmented exomes perform far better for known disease causing genes, but less well for other genes and in untranslated regions. Increasing whole genome coverage improves this discrepancy with an average coverage of {small tilde}45x representing the cross over point where performance equals that of exome capture for disease causing genes. A combination of some genome-wide coverage and augmented exon coverage may offer the most cost effective solution for clinical grade genome sequencing today. In summary, this coverage metric provides transparency regarding the current state of next-generation sequencing for clinical medicine and will inform genotype interpretation, technology improvement, and sequencing platform choices for physicians and laboratories. We provide an application on precision.fda.gov (Coverage of Key Genes app) to calculate this metric.

Genomics

Elusive Plasmodium Species Complete the Human Malaria Genome Set

Despite the huge international endeavor to understand the genomic basis of malaria biology, there remains a lack of information about two human-infective species: Plasmodium malariae and P. ovale. The former is prevalent across all malaria endemic regions and able to recrudesce decades after the initial infection. The latter is a dormant stage hypnozoite-forming species, similar to P. vivax. Here we present the newly assembled reference genomes of both species, thereby completing the set of all human-infective Plasmodium species. We show that the P. malariae genome is markedly different to other Plasmodium genomes and relate this to its unique biology. Using additional draft genome assemblies, we confirm that P. ovale consists of two cryptic species that may have diverged millions of years ago. These genome sequences provide a useful resource to study the genetic basis of human-infectivity in Plasmodium species.

Genomics