Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Searching and Indexing Genomic Databases via Kernelization

The rapid advance of DNA sequencing technologies has yielded databases of thousands of genomes. To search and index these databases effectively, it is important that we take advantage of the similarity between those genomes. Several authors have recently suggested searching or indexing only one reference genome and the parts of the other genomes where they differ. In this paper we survey the twenty-year history of this idea and discuss its relation to kernelization in parameterized complexity.

Bioinformatics

FIQT: a simple, powerful method to accurately estimate effect sizes in genome scans

Genome scans, including both genome-wide association studies and deep sequencing, continue to discover a growing number of significant association signals for various traits. However, often variants meeting genome-wide significance criteria explain far less of the overall trait variance than \"sub-threshold\" association signals. To extract these sub-threshold signals, there is a need for methods which accurately estimate the mean of all (normally-distributed) test-statistics from a genome scan (i.e., Z-scores). This is currently achieved by the difficult procedures of adjusting all Z-score [Formula] statistics for \"winners curse\" (multiple testing). Given that multiple testing adjustments are much simpler for p-values, we propose a method for estimating Z-scores means by i) first adjusting their p-values for multiple testing and then ii) transforming the adjusted p-values to upper tail Z-scores with the sign of the original statistics. Because a False Discovery Rate (FDR) procedure is used for multiple testing adjustment, we denote this method FDR Inverse Quantile Transformation (FIQT). When compared to competitors, e.g. Empirical Bayes (including proposed improvements), FIQT is more i) accurate and ii) computationally efficient by orders of magnitude. Its accuracy advantage is substantial at larger sample sizes and/or moderate numbers of association signals. Practical application of FIQT to Z-scores from the first Psychiatric Genetic Consortium (PGC) schizophrenia predicts a non-trivial fraction of the significant signal regions from the subsequent published PGC schizophrenia studies. Finally, we suggest that FIQT might be i) used to improve subject level risk prediction and ii) further improved by modelling the noncentrality of [Formula] statistics.

Genetics

Roary: Rapid large-scale prokaryote pan genome analysis

SummaryA typical prokaryote population sequencing study can now consist of hundreds or thousands of isolates. Interrogating these datasets can provide detailed insights into the genetic structure of of prokaryotic genomes. We introduce Roary, a tool that rapidly builds large-scale pan genomes, identifying the core and dispensable accessory genes. Roary makes construction of the pan genome of thousands of prokaryote samples possible on a standard desktop without compromising on the accuracy of results. Using a single CPU Roary can produce a pan genome consisting of 1000 isolates in 4.5 hours using 13 GB of RAM, with further speedups possible using multiple processors.\n\nAvailability and implementationRoary is implemented in Perl and is freely available under an open source GPLv3 license from http://sanger-pathogens.github.io/Roary\n\nContactroary@sanger.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Signatures of Dobzhansky-Muller Incompatibilities in the Genomes of Recombinant Inbred Lines

In the construction of Recombinant Inbred Lines (RILs) from two divergent inbred parents certain genotype (or epigenotype) combinations may be functionally \"incompatible\" when brought together in the genomes of the progeny, thus resulting in sterility or lower fertility. Natural selection against these epistatic combinations during inbreeding can change haplotype frequencies and distort linkage disequilibrium (LD) relations between loci within and across chromosomes. These LD distortions have received increased experimental attention, because they point to genomic regions that may drive Dobzhansky-Muller-type of reproductive isolation and, ultimately, speciation in the wild. Here we study the selection signatures of two-locus epistatic incompatibility models and quantify their impact on the genetic composition of the genomes of 2-way RILs obtained by selfing. We also consider the biases introduced by breeders when trying to counteract the loss of lines by selectively propagating only viable seeds. Building on our theoretical results, we develop model-based maximum likelihood (ML) tests which can be employed in pairwise genome scans for incompatibility loci using multi-locus genotype data. We illustrate this ML approach in the context of two published A.thaliana RIL panels. Our work lays the theoretical foundation for studying more complex systems such as RILs obtained by sibling mating and/or from multi-parental crosses.

Genetics

Correcting bias from stochastic insert size in read pair data — applications to structural variation detection and genome assembly

Insert size distributions from paired read protocols are used for inference in bioinformatic applications such as genome assembly and structural variation detection. However, many of the models that are being used are subject to bias. This bias arises when we assume that all insert sizes within a distribution are equally likely to be observed, when in fact, size matters. These systematic errors exist in popular software even when the assumptions made about data are true. We have previously shown that bias occurs for scaffolders in genome assembly. Here, we generalize the theory and demonstrate that it is applicable in other contexts. We provide examples of bias in state-of the-art software and improve them using our model. One key application of our theory is structural variation detection using read pairs. We show that an incorrect null-hypothesis is commonly used in popular tools and can be corrected using our theory. Furthermore, we approximate the smallest size of indels that are possible to discover given an insert size distribution. Two other applications are inference of insert size distribution on \emph{de novo} genome assemblies and error correction of genome assemblies using mated reads. Our theory is implemented in a tool called GetDistr (\url{https://github.com/ksahlin/GetDistr}).

Bioinformatics

Evolutionary quantitative genomics of Populus trichocarpa

Forest trees generally show high levels of local adaptation and efforts focusing on understanding adaptation to climate will be crucial for species survival and management.\n\nMerging quantitative genetics and population genomics, we studied the molecular basis of climate adaptation in 433 Populus trichocarpa (black cottonwood) genotypes originating across western North America. Variation in 74 field-assessed traits (growth, ecophysiology, phenology, leaf stomata, wood, and disease resistance) was investigated for signatures of selection (comparing QST-FST) using clustering of individuals by climate of origin. 29,354 SNPs were investigated employing three different outlier detection methods.\n\nNarrow-sense QST for 53% of distinct field QST traits was significantly divergent from expectations of neutrality (indicating adaptive trait variation); 2,855 SNPs showed signals of diversifying selection and of these, 118 SNPs (within 81 genes) were associated with adaptive traits (based on significant QST). Many SNPs were putatively pleiotropic for functionally uncorrelated adaptive traits, such as autumn phenology, height, and disease resistance.\n\nEvolutionary quantitative genomics in P. trichocarpa provides an enhanced understanding regarding the molecular basis of climate-driven selection in forest trees. We highlight that important loci underlying adaptive trait variation also show relationship to climate of origin.\n\nAuthor summaryComparisons between population differentiation on the basis of quantitative traits and neutral genetic markers inform about the importance of natural selection, genetic drift and gene flow for local adaptation of populations. Here, we address fundamental questions regarding the molecular basis of adaptation in undomesticated forest tree populations to past climatic environments by employing an integrative quantitative genetics and landscape genomics approach. Marker-inferred relatedness was estimated to obtain the narrow-sense estimate of population differentiation in wild populations. We analyzed an unstructured population of common garden grown Populus trichocarpa individuals to uncover different extents of variation for a suite of field traits, wood quality and pathogen resistance with temperature and precipitation. We consider our approach the most comprehensive, as it uncovers the molecular mechanisms of adaptation using multiple methods and tests. We provide a detailed outline of the required analyses for studying adaptation to the environment in a population genomics context to better understand the species potential adaptive capacity to future climatic scenarios.

Evolutionary Biology

Complete assembly of novel environmental bacterial genomes by MinIONTM sequencing

In this study, we adapt a protocol for the growth of previously uncultured environmental bacterial isolates, to make it compatible with whole genome sequencing. We demonstrate that in combination with the MinION sequencing device, complete assemblies can be derived, allowing genomic comparisons to be made. This approach allows rapid, inexpensive and straightforward discovery, and genomic analysis, of previously uncultured prokaryotic genomes, and brings greater ownership of all parts of the sequencing process back to individual researchers.

Microbiology

Resolving Complex Structural Genomic Rearrangements using a Randomized Approach

Complex chromosomal rearrangements consist of structural genomic alterations involving multiple instances of deletions, duplications, inversions, or translocations that co-occur either on the same chromosome or represent different overlapping events on homologous chromosomes. We present SVelter, an algorithm that first identifies regions of the genome suspected to harbor a complex event and then iteratively rearranges the local genome structure, in a randomized fashion, with each structure scored against characteristics of the observed sequencing data. We show that SVelter is able to accurately reconstruct these regions when compared to well-characterized genomes that have been deep sequenced with both short and long read technologies.

Bioinformatics

Distinct genomic and epigenomic features demarcate hypomethylated blocks in colon cancer

Background. Large mega base-pair genomic regions show robust alterations in DNA methylation levels in multiple cancers, a vast majority of which are hypo-methylated in cancers. These regions are generally bounded by CpG islands, overlap with Lamin Associated Domains and Large organized chromatin lysine modifications, and are associated with stochastic variability in gene expression. Given the size and consistency of hypo-methylated blocks (HMB) across cancer types, their immediate causes are likely to be encoded in the genomic region near HMB boundaries, in terms of specific genomic or epigenomic signatures. However, a detailed characterization of the HMB boundaries has not been reported.\n\nMethod. Here, we focused on ~13k HMBs, encompassing approximately half the genome, identified in colon cancer. We analyzed a number of distinguishing features at the HMB boundaries including transcription factor (TF) binding motifs, various epigenomic marks, and chromatin structural features.\n\nResult. We found that the classical promoter epigenomic mark - H3K4me3, is highly enriched at HMB boundaries, as are CTCF bound sites. HMB boundaries harbor distinct combinations of TF motifs. Our Random Forest model based on TF motifs can accurately distinguish boundaries not only from regions inside and outside HMBs, but surprisingly, from active promoters as well. Interestingly, the distinguishing TFs and their interacting proteins are involved in chromatin modification. Finally, HMB boundaries significantly coincide with the boundaries of Topologically Associating Domains of the chromatin.\n\nConclusion. Our analyses suggest that the overall architecture of HMBs is guided by pre-existing chromatin architecture, and are associated with aberrant activity of promoter-like sequences at the boundary.

Bioinformatics

Identification and analysis of integrons and cassette arrays in bacterial genomes

Integrons recombine gene arrays and favor the spread of antibiotic resistance. Their broader roles in bacterial adaptation remain mysterious, partly due to lack of computational tools. We made a program - IntegronFinder - to identify integrons with high accuracy and sensitivity. IntegronFinder is available as a standalone program and as a web application. It searches for attC sites using covariance models, for integron-integrases using HMM profiles, and for other features (promoters, attl site) using pattern matching. We searched for integrons, integron-integrases lacking attC sites, and clusters of attC sites lacking a neighboring integron-integrase in bacterial genomes. All these elements are especially frequent in genomes of intermediate size. They are missing in some key phyla, such as -Proteobacteria, which might reflect selection against cell lineages that acquire integrons. The similarity between attC sites is proportional to the number of cassettes in the integron, and is particularly low in clusters of attC sites lacking integron-integrases. The latter are unexpectedly abundant in genomes lacking integron-integrases or their remains, and have a large novel pool of cassettes lacking homologs in the databases. They might represent an evolutionary step between the acquisition of genes within integrons and their stabilization in the new genome.

Bioinformatics

Whole genome sequencing of 56 Mimulus individuals illustrates population structure and local selection

Across western North America, Mimulus guttatus exists as many local populations adapted to site-specific challenges including salt spray, temperature, water availability, and soil chemistry. Gene flow between locally adapted populations will effect genetic diversity in both local demes and across the larger meta-population. A single population of annual M. guttatus from Iron Mountain, Oregon (IM) has been extensively studied and we here building off this research by analyzing whole genome sequences from 34 inbred lines from IM in conjunction with sequences from 22 Mimulus individuals from across the geographic range. Three striking features of these data address hypotheses about migration and selection in a locally adapted population. First, we find very high intra-population polymorphism (synonymous {pi} = 0.033). Variation outside genes may be even higher, but is difficult to estimate because excessive divergence affects read mapping. Second, IM exhibits a significantly positive genome-wide average for Tajimas D. This indicates allele frequencies are typically more intermediate than expected from neutrality, opposite the pattern observed in other species. Third, IM exhibits a distinctive haplotype structure. There is a genome-wide excess of positive associations between minor alleles; consistent with an important effect of gene flow from nearby Mimulus populations. The combination of multiple data types, including a novel, tree-based analytic method and estimates for structural polymorphism (inversions) from previous genetic mapping studies, illustrates how the balance of strong local selection, limited dispersal, and meta-population dynamics manifests across the genome.

Evolutionary Biology

Unearthing new genomic markers of drug response by improved measurement of discriminative power

BackgroundOncology drugs are only effective in a small proportion of cancer patients. Our current ability to identify these responsive patients before treatment is still poor in most cases. Thus, there is a pressing need to discover response markers for marketed and research oncology drugs in order to improve patient survival, reduce healthcare costs and enhance success rates in clinical trials. Screening these drugs against a large panel of cancer cell lines has been employed to discover new genomic markers of in vitro drug response, which can now be further evaluated on more accurate tumour models. However, while the identification of discriminative markers among thousands of candidate drug-gene associations in the data is error-prone, an appraisal of the effectiveness of such detection task is currently lacking.\n\nResultsHere we present a new non-parametric method to measuring the discriminative power of a drug-gene association. This is enabled by the identification of an auxiliary threshold posing this task as a binary classification problem. Unlike parametric statistical tests, the adopted non-parametric test has the advantage of not making strong assumptions about the data distorting the identification of genomic markers. Furthermore, we introduce a new benchmark to further validate these markers in vitro using more recent data not used to identify the markers. The application of this new methodology has led to the identification of 128 new genomic markers distributed across 61% of the analysed drugs, including 5 drugs without previously known markers, which were missed by the MANOVA test initially applied to analyse data from the Genomics of Drug Sensitivity in Cancer consortium.\n\nAbbreviation

Bioinformatics

The role of deleterious substitutions in crop genomes

Populations continually incur new mutations with fitness effects ranging from lethal to adaptive. While the distribution of fitness effects (DFE) of new mutations is not directly observable, many mutations likely have either no effect on organismal fitness or are deleterious. Historically, it has been hypothesized that a population may carry many mildly deleterious variants as segregating variation, which reduces the mean absolute fitness of the population. Recent advances in sequencing technology and sequence conservation-based metrics for inferring the functional effect of a variant permit examination of the persistence of deleterious variants in populations. The issue of segregating deleterious variation is particularly important for crop improvement, because the demographic history of domestication and breeding allows deleterious variants to persist and reach moderate frequency, potentially reducing crop productivity. In this study, we use exome resequencing of fifteen barley accessions and genome resequencing of eight soybean accessions to investigate the prevalence of deleterious SNPs in the protein-coding regions of the genomes of two crops. We conclude that individual cultivars carry hundreds of deleterious SNPs on average, and that nonsense variants make up a minority of deleterious SNPs. Our approach annotates known phenotype-altering variants as deleterious more frequently than the genome-wide average, suggesting that putatively deleterious variants are likely to affect phenotypic variation. We also report the implementation of a SNP annotation tool (BAD_Mutations) that makes use of a likelihood ratio test based on alignment of all currently publicly available Angiosperm genomes.

Genetics

Revolutionising Public Health Reference Microbiology using Whole Genome Sequencing: Salmonella as an exemplar

Advances in whole genome sequencing (WGS) platforms and DNA library preparation have led to the development of methods for high throughput sequencing of bacterial genomes at a relatively low cost (Loman et al. 2012; Medini et al. 2008). WGS offers unprecedented resolution for determining degrees of relatedness between strains of bacterial pathogens and has proven a powerful tool for microbial population studies and epidemiological investigations (Harris et al. 2010; Lienau et al. 2011; Holt et al. 2009; Ashton, Peters, et al. 2015). The potential utility of WGS to public health microbiology has been highlighted previously (Koser et al. 2012; Kwong et al. 2013; Reuter et al. 2013; Joensen et al. 2014; Nair et al. 2014; Bakker et al. 2014; DAuria et al. 2014). Here we report, for the first time, the routine use of WGS as the primary test for identification, surveillance and outbreak investigation by a national reference laboratory. We present data on how this has revolutionised public health microbiology for one of the most common bacterial pathogens in the United Kingdom, the Salmonellae.\n\nDATA SUMMARY1. PHE Salmonella sequencing data is deposited in the Sequence Read Archive in BioProject PRJNA248792.\n\nIMPACT STATEMENTThe first human genome cost around $3 billion, and took around 10 years to complete. Advances in DNA sequencing technology (also referred to as whole genome sequencing (WGS)) allow the same feat to be accomplished today for less than $10000 and less than 2 weeks. This remarkable improvement in technology has also led to a step change in microbiology, increasing our understanding of the evolution of major human pathogens such as Yersinia pestis, Salmonella Typhi and Mycobacterium tuberculosis. While these kinds of academic studies provide unparalleled context for public health action, until now, this approach has not been routinely employed at the frontline. At Public Health England, WGS has been implemented for routine public health identification, characterisation and typing of an important human pathogen, Salmonella, replacing methods that have changed little over the last 100 years. Analysis of WGS data has identified outbreaks that were previously undetectable and been used to infer rare antimicrobial resistance patterns. This paper will serve as a notification to the community of the methods PHE are using, and will be of great use to other public health labs considering switching to WGS.

Microbiology

Practical Guidelines for Secure Cloud Computing using Genomic Data

Cloud security challenges Cloud security challenges Cloud security guidelines Security and Privacy Publications: Large scale genomics studies involving thousands of whole genome or exome sequences are underway1 on Cloud. While Cloud provides many conveniences for genomics research, it also raises concerns regarding large scale hacking, bad press, potential loss of patient privacy and the resulting loss of patient trust. Cloud providers argue that they have significant investments and expertise in security and, therefore, Cloud is equally secure, if not more so, compared to on-premise infrastructure. This gap in assessment of Cloud security is, in part, due to a fast evolving and largely unfamiliar technology stack for genomics data owners.\n\nWhat makes the Cloud security landscape discussion challenging is that security reco ...

Bioinformatics

Whole-genome sequencing uncovers cryptic and hybrid species among Atlantic and Pacific cod-fish

Speciation often involves the splitting of a lineage and the adaptation of daughter lineages to different environments. It may also involve the merging of divergent lineages, thus creating a stable homoploid hybrid species1 that constructs a new ecological niche by transgressing2 the ecology of the parental types. Hybrid speciation may also contribute to enigmatic and cryptic biodiversity in the sea.3,4 The enigmatic walleye pollock, which is not a pollock at all but an Atlantic cod that invaded the Pacific 3.8 Mya,5 differs considerably from its presumed closest relatives, the Pacific and Atlantic cod. Among the Atlantic cod, shallow-water coastal and deep-water migratory frontal ecotypes are associated with highly divergent genomic islands;6,7 however, intermediates remain an enigma.8 Here, we performed whole-genome sequencing of over 200 individuals using up to 33 million SNPs based on genotype likelihoods9 and showed that the evolutionary status of walleye pollock is a hybrid species: it is a hybrid between Arctic cod and Atlantic cod that transgresses the ecology of its parents. For the first time, we provide decisive evidence that the Atlantic cod coastal and frontal ecotypes are separate species that hybridized, leading to a true-breeding hybrid species that differs ecologically from its parents. We refute monophyly and dichotomous branching of these taxa, and stress the importance of looking beyond branching trees at admixture and hybridity. Our study demonstrates the power of whole-genome sequencing and population genomics in providing deep insights into fundamental processes of speciation. Our study was a starting point for further work aimed at examining the criteria of hybrid speciation,10 selection, sterility and structural chromosomal variation11 among cod-fish, which are among the most important fish stocks in the world. The hybrid nature of both the walleye pollock and Atlantic cod raises the question concerning the extent to which very profitable fisheries12,13 depend on hybrid vigour. Our results have implications for management of marine resources in times of rapid climate change.14,15

Preprint

Inferring population size history from large samples of genome wide molecular data - an approximate Bayesian computation approach

Inferring the ancestral dynamics of effective population size is a long-standing question in population genetics, which can now be tackled much more accurately thanks to the massive genomic data available in many species. Several promising methods that take advantage of whole-genome sequences have been recently developed in this context. However, they can only be applied to rather small samples, which limits their ability to estimate recent population size history. Besides, they can be very sensitive to sequencing or phasing errors. Here we introduce a new approximate Bayesian computation approach named PopSizeABC that allows estimating the evolution of the effective population size through time, using a large sample of complete genomes. This sample is summarized using the folded allele frequency spectrum and the average zygotic linkage disequilibrium at different bins of physical distance, two classes of statistics that are widely used in population genetics and can be easily computed from unphased and unpolarized SNP data. Our approach provides accurate estimations of past population sizes, from the very first generations before present back to the expected time to the most recent common ancestor of the sample, as shown by simulations under a wide range of demographic scenarios. When applied to samples of 15 or 25 complete genomes in four cattle breeds (Angus, Fleckvieh, Holstein and Jersey), PopSizeABC revealed a series of population declines, related to historical events such as domestication or modern breed creation. We further highlight that our approach is robust to sequencing errors, provided summary statistics are computed from SNPs with common alleles.

Genetics

Increasing the Efficiency of Genome-wide Association Mapping via Hidden Markov Models

With the rapid production of high dimensional genetic data, one major challenge in genome-wide association studies is to develop effective and efficient statistical tools to resolve the low power problem of detecting causal SNPs with low to moderate susceptibility, whose effects are often obscured by substantial background noises. Here we present a novel method that serves as an optimal technique for reducing background noises and improving detection power in genome-wide association studies. The approach uses hidden Markov model and its derivate Markov hidden Markov model to estimate the posterior probabilities of a markers being in an associated state. We conducted extensive simulations based on the human whole genome genotype data from the GlaxoSmithKline-POPRES project to calibrate the sensitivity and specificity of our method and compared with many popular approaches for detecting positive signals including the{chi} 2 test for association and the Cochran-Armitage trend test. Our simulation results suggested that at very low false positive rates (< 10-6), our method reaches the power of 0.9, and is more powerful than any other approaches, when the allelic effect of the causal variant is non-additive or unknown. Application of our method to the data set generated by Welcome Trust Case Control Consortium using 14,000 cases and 3,000 controls confirmed its powerfulness and efficiency under the context of the large-scale genome-wide association studies.

Genetics