Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

Haplotype synthesis analysis in public reference data reveals functional variants underlying known genome-wide associated susceptibility loci

The functional mechanisms underlying disease association identified by Genome-wide Association Studies remain unknown for susceptibility loci located outside gene coding regions. In addition to the regulation of gene expression, synthesis of effects from multiple surrounding functional variants has been suggested as an explanation of hard-to-interpret associations.\n\nHere, we define filter criteria based on linkage disequilibrium measures and allele frequencies which reflect expected properties of synthesizing variant sets. For eligible candidate sets we search for those haplotypes that are highly correlated with the risk alleles of a genome-wide associated variant.\n\nWe applied our methods to 1,000 Genomes reference data and confirmed Crohns Disease and Type 2 Diabetes susceptibility loci. Of these, a proportion of 32% allowed explanation by three-variant-haplotypes carrying at least two functional variants, as compared to a proportion of 16% for random variants (P = 2.92 {middle dot} 10-6). More importantly, we detected examples of known loci whose association can fully be explained by surrounding missense variants: three missense variants from MUC19 synthesize rs11564258 (L0C105369736/MUC19, intron; Crohns Disease). Next, rs2797685 (PER3, intron; Crohns Disease) is synthesized by a 57 kilobase haplotype defined by five missense variants from PER3 and three missense variants from UTS2. Finally, the association of rs7178572 (HMG20A, intron; Type 2 Diabetes) can be explained by the synthesis of eight haplotypes, each carrying at least one missense variant in either PEAK1, TBC1D2B, CHRNA5 or ADAMTS7.\n\nIn summary, application of our new methods highlights the potential of synthesis analysis to guide functional follow-up investigation of findings from association studies.

Bioinformatics

FINEMAP: Efficient variable selection using summary data from genome-wide association studies

MotivationThe goal of fine-mapping in genomic regions associated with complex diseases and traits is to identify causal variants that point to molecular mechanisms behind the associations. Recent fine-mapping methods using summary data from genome-wide association studies rely on exhaustive search through all possible causal configurations, which is computationally expensive.\n\nResultsWe introduce FINEMAP, a software package to efficiently explore a set of the most important causal configurations of the region via a shotgun stochastic search algorithm. We show that FINEMAP produces accurate results in a fraction of processing time of existing approaches and is therefore a promising tool for analyzing growing amounts of data produced in genome-wide association studies.\n\nAvailabilityFINEMAP v1.0 is freely available for Mac OS X and Linux at http://www.christianbenner.com.\n\nContact: christian.benner@helsinki.fi, matti.pirinen@helsinki.fi

Genetics

One Codex: A Sensitive and Accurate Data Platform for Genomic Microbial Identification

High-throughput sequencing (HTS) is increasingly being used for broad applications of microbial characterization, such as microbial ecology, clinical diagnosis, and outbreak epidemiology. However, the analytical task of comparing short sequence reads against the known diversity of microbial life has proved to be computationally challenging. The One Codex data platform was created with the dual goals of analyzing microbial data against the largest possible collection of microbial reference genomes, as well as presenting those results in a format that is consumable by applied end-users. One Codex identifies microbial sequences using a \"k-mer based\" taxonomic classification algorithm through a web-based data platform, using a reference database that currently includes approximately 40,000 bacterial, viral, fungal, and protozoan genomes. In order to evaluate whether this classification method and associated database provided quantitatively different performance for microbial identification, we created a large and diverse evaluation dataset containing 50 million reads from 10,639 genomes, as well as sequences from six organisms novel species not be included in the reference databases of any of the tested classifiers. Quantitative evaluation of several published microbial detection methods shows that One Codex has the highest degree of sensitivity and specificity (AUC = 0.97, compared to 0.82-0.88 for other methods), both when detecting well-characterized species as well as newly sequenced, \"taxonomically novel\" organisms.

Bioinformatics

Comparing cancer cell lines and tumor samples by genomic profiles

Cancer cell lines are often used in laboratory experiments as models of tumors, although they can have substantially different genetic and epigenetic profiles compared to tumors. We have developed a general computational method - TumorComparer - to systematically quantify similarities and differences between tumor material when detailed genetic and molecular profiles are available. The comparisons can be flexibly tailored to a particular biological question by placing a higher weight on functional alterations of interest ( weighted similarity). In a first pan-cancer application, we have compared 260 cell lines from the Cancer Cell Line Encyclopaedia (CCLE) and 1914 tumors of six different cancer types from The Cancer Genome Atlas (TCGA), using weights to emphasize genomic alterations that frequently recur in tumors. We report the potential suitability of particular cell lines as tumor models and identify apparently unsuitable outlier cell lines, some of which are in wide use, for each of the six cancer types. In future, this weighted similarity method may be generalized for use in a clinical setting to compare patient profiles consisting of genomic patterns combined with clinical attributes, such as diagnosis, treatment and response to therapy.

Cancer Biology

The Great Migration and African-American genomic diversity

Genetic studies of African-Americans identify functional variants, elucidate historical and genealogical mysteries, and reveal basic biology. However, African-Americans have been under-represented in genetic studies, and little is known about nation-wide patterns of genomic diversity in the population. Here, we present a comprehensive assessment of African-American genomic diversity using genotype data from nationally and regionally representative cohorts. We find higher African ancestry in southern United States compared to the North and West. We show that relatedness patterns track north- and west-bound routes followed during the Great Migration, suggesting that admixture occurred predominantly in the South prior to the Civil War and that ancestry-biased migration is responsible for regional differences in ancestry. Rare genetic traits among African-Americans can therefore be shared over long geographic distances along the Great Migration routes, yet their distribution over short distances remains highly structured. This study clarifies the role of recent demography in shaping African-American genomic diversity.

Preprint

Whole-genome modeling accurately predicts quantitative traits, as revealed in plants.

Many adaptive events in natural populations, as well as response to artificial selection, are caused by polygenic action. Under selective pressure, the adaptive traits can quickly respond via small allele frequency shifts spread across numerous loci. We hypothesize that a large proportion of current phenotypic variation between individuals may be best explained by population admixture.\n\nWe thus consider the complete, genome-wide universe of genetic variability, spread across several ancestral populations originally separated. We experimentally confirmed this hypothesis by predicting the differences in quantitative disease resistance levels among accessions in the wild legume Medicago truncatula. We discovered also that variation in genome admixture proportion explains most of phenotypic variation for several quantitative functional traits, but not for symbiotic nitrogen fixation. We shown that positive selection at the species level might not explain current, rapid adaptation.\n\nThese findings prove the infinitesimal model as a mechanism for adaptation of quantitative phenotypes. Our study produced the first evidence that the whole-genome modeling of DNA variants is the best approach to describe an inherited quantitative trait in a higher eukaryote organism and proved the high potential of admixture-based analyses. This insight contribute to the understanding of polygenic adaptation, and can accelerate plant and animal breeding, and biomedicine research programs.

Genetics

Whole genome duplication in coast redwood (Sequoia sempervirens) and its implications for explaining the rarity of polyploidy in conifers

SO_SCPLOWUMMARYC_SCPLOWO_LIWhereas polyploidy is common and an important evolutionary factor in most land plant lineages it is a real rarity in gymnosperms. Coast redwood (Sequoia sempervirens) is the only hexaploid conifer and one of just two naturally polyploid conifer species. Numerous hypotheses about the mechanism of polyploidy in Sequoia and parental genome donors have been proffered over the years, primarily based on morphological and cytological data, but it remains unclear how Sequoia became polyploid and why this lineage overcame an apparent gymnosperm barrier to whole-genome duplication (WGD).\nC_LIO_LIWe sequenced transcriptomes and used phylogenetic inference, Bayesian concordance analysis, and paralog age distributions to resolve relationships among gene copies in hexaploid coast redwood and its close relatives.\nC_LIO_LIOur data show that hexaploidy in the coast redwood lineage is best explained by autopolyploidy or, if there was allopolyploidy, this was restricted to within the Californian redwood clade. We found that duplicate genes have more similar sequences than would be expected given evidence from fossil guard cell size which suggest that polyploidy dates to the Eocene.\nC_LIO_LIConflict between molecular and fossil estimates of WGD can be explained if diploidization occurred very slowly following whole genome duplication. We extrapolate from this to suggest that the rarity of polyploidy in conifers may be due to slow rates of diploidization in this clade.\nC_LI

Plant Biology

CRISPResso: sequencing analysis toolbox for CRISPR genome editing

To the Editor To the Editor References Recent progress in genome editing technologies, in particular the CRISPR-Cas9 nuclease system, has provided new opportunities to investigate the biological functions of genomic sequences by targeted mutagenesis [1-4]. Briefly, Cas9 may be directed by a chimeric single guide RNA (sgRNA) to a target genomic sequence upstream of a protospacer adjacent motif (PAM) for cleavage. Double strand breaks (DSBs) resulting from site-specific Cas9 cleavage can be resolved by endogenous DNA repair pathways such as non-homologous end joining (NHEJ) or homology-directed repair (HDR). These repair mechanisms result in a spectrum of diverse outcomes including insertions, deletions, nucleotide substitutions, and, in the case of HDR, recombination of extrachromosomal donor sequences [1-3, 5 ...

Bioinformatics

AGOUTI: improving genome assembly and annotation using transcriptome data

SummaryCurrent genome assemblies consist of thousands of contigs. These incomplete and fragmented assemblies lead to errors in gene identification, such that single genes spread across multiple contigs are annotated as separate gene models. We present AGOUTI (Annotated Genome Optimization Using Transcriptome Information), a tool that uses RNA-seq data to simultaneously combine contigs into scaffolds and fragmented gene models into single models. We show that AGOUTI improves both the contiguity of genome assemblies and the accuracy of gene annotation, providing updated versions of each as output.\n\nAvailabilityThe software is implemented in python and is available from github.com/svm-zhang/AGOUTI.\n\nContactsimozhan@indiana.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Interactive Analytics for Very Large Scale Genomic Data

Large scale genomic sequencing is now widely used to decipher questions in diverse realms such as biological function, human diseases, evolution, ecosystems, and agriculture. With the quantity and diversity these data harbor, a robust and scalable data handling and analysis solution is desired. Here we present interactive analytics using public cloud infrastructure and distributed computing database Dremel and developed according to the standards of Global Alliance for Genomics and Health, to perform information compression, comprehensive quality controls, and biological information retrieval in large volumes of genomic data. We demonstrate that such computing paradigms can provide orders of magnitude faster turnaround for common analyses, transforming long-running batch jobs submitted via a Linux shell into questions that can be asked from a web browser in seconds.

Preprint

Comparative evaluation of the genomes of common bacterial members of the Drosophila intestinal community

Drosophila melanogaster is an excellent model to explore the molecular exchanges that occur between an animal intestine and their microbial passengers. For example, groundbreaking studies in flies uncovered a sophisticated web of host responses to intestinal bacteria. The outcomes of these responses define critical events in the host, such as the establishment of immune responses, access to nutrients, and the rate of larval development. Despite our steady march towards illuminating the host machinery that responds to bacterial presence in the gut, we know remarkably little about the microbial products that influence bacterial association with a fly host. To address this deficiency, we sequenced and characterized the genomes of three common Drosophila-associated microbes: Lactobacillus plantarum, Lactobacillus brevis and Acetobacter pasteurianus. In each case, we compared the genomes of Drosophila-associated strains to the genomes of strains isolated from alternative sources. This approach allowed us to identify molecular functions common to Drosophila-associated microbes, and, in the case of A. pasteurianus, to identify genes that are essential for association with the host. Of note, many of the gene products unique to fly-associated strains have established roles in the stabilization of host-microbe interactions. We believe that these data provide a valuable starting point for a more thorough examination of the microbial perspective on host-microbe relationships.

Microbiology

Mapping challenging mutations by whole-genome sequencing

Whole-genome sequencing provides a rapid and powerful method for identifying mutations on a global scale, and has spurred a renewed enthusiasm for classical genetic screens in model organisms. The most commonly characterized category of mutation consists of monogenic, recessive traits, due to their genetic tractability. Therefore, most of the mapping methods for mutation identification by whole-genome sequencing are directed toward alleles that fulfill those criteria (i.e., single-gene, homozygous variants). However, such approaches are not entirely suitable for the characterization of a variety of more challenging mutations, such as dominant and semi-dominant alleles or multigenic traits. Therefore, we have developed strategies for the identification of those classes of mutations, using polymorphism mapping in Caenorhabditis elegans as our model for validation. We also report an alternative approach for mutation identification from traditional recombinant crosses, and a solution to the technical challenge of sequencing sterile or terminally arrested strains where population size is limiting. The methods described herein extend the applicability of whole-genome sequencing to a broader spectrum of mutations, including classes that are difficult to map by traditional means.

Genetics

Tempo and mode of genome evolution in a 50,000-generation experiment

Adaptation depends on the rates, effects, and interactions of many mutations. We analyzed 264 genomes from 12 Escherichia coli populations to characterize their dynamics over 50,0 generations. The trajectories for genome evolution in populations that retained the ancestral mutation rate fit a model where most fixed mutations are beneficial, the fraction of beneficial mutations declines as fitness rises, and neutral mutations accumulate at a constant rate. We also compared these populations to lines evolved under a mutation-accumulation regime that minimizes selection. Nonsynonymous mutations, intergenic mutations, insertions, and deletions are overrepresented in the long-term populations, supporting the inference that most fixed mutations are favored by selection. These results illuminate the shifting balance of forces that govern genome evolution in populations adapting to a new environment.

Evolutionary Biology

Using Genome Wide Estimates of Heritability to Examine the Relevance of Gene-Environment Interplay

We use genome-wide data from the third generation respondents of the Framing-ham Heart Study to estimate heritability in body mass index using different quantities of the measured genotype. Heritability decreases rapidly when SNPs implicated by a genome-wide association study are removed but shows essentially no decline when SNPs implicated by a gene-environment interaction in a second genome-wide analysis are removed. This second result is highlighted by our additional finding that the SNPs which explain heritability amongst a subsample defined by higher educational attainment explain no heritability of the heritability in the lower education group, and vice-versa. Finally, we do find consistent heritability estimates when we compare family-based estimates versus those based on measured genotype.

Genetics

Phenotypic plasticity promotes balanced polymorphism in periodic environments by a genomic storage effect

Phenotypic plasticity is known to evolve in perturbed habitats, where it alleviates the deleterious effects of selection. But the effects of plasticity on levels of genetic polymorphism, an important precursor to adaptation in temporally varying environments, are unclear. Here we develop a haploid, two-locus population-genetic model to describe the interplay between a plasticity modifier locus and a target locus subject to periodically varying selection. We find that the interplay between these two loci can produce a \"genomic storage effect\" that promotes balanced polymorphism over a large range of parameters, in the absence of all other conditions known to maintain genetic variation. The genomic storage effect arises as recombination allows alleles at the two loci to escape more harmful genetic backgrounds and associate in haplotypes that persist until environmental conditions change. Using both Monte Carlo simulations and analytical approximations we quantify the strength of the genomic storage effect across a range of selection pressures, recombination rates, plasticity modifier effect sizes, and environmental periods.

Evolutionary Biology

Implementation of an Open Source Software solution for Laboratory Information Management and automated RNAseq data analysis in a large-scale Cancer Genomics initiative using BASE with extension package Reggie.

BackgroundLarge-scale cancer genomics initiatives and next-generation sequencing for transcriptome profiling allow for detailed molecular characterization of tumors, and provide opportunities for clinical tools to improve diagnosis, prognosis, and treatment decisions. Laboratory information, data management, and data sharing in large-scale genomics projects is a challenge. Aiming to introduce such technologies in a clinical setting offer additional challenges associated with requirements of short lead-times and specialized tracking of biomaterials, data, and analysis results.\n\nResultsUsing the free open-source BioArray Software Environment (BASE) and extension package Reggie we have implemented a laboratory information management system and an automated RNAseq data analysis pipeline that successfully manage a large regional cancer genomics initiative. The system manages enrolled cancer patients, tumor biopsies, extraction of nucleic acid, and whole transcriptome RNA-sequencing through to data analysis and quality control. The implementation offers integration of laboratory equipment and operating procedures, and information tracking in a module based fashion enabling efficient and flexible use of personnel resources. The system provides two-factor authentication and transaction control and seamless integration of freely available software for RNAseq analysis such as Tophat, Cufflinks, and Picard. As of February 2016 more than 8000 patients and over 6000 tumor biopsies have been successfully processed. Lead-time from biopsy arrival to summarized reports based on RNAseq data is less than 5 days, in line with regional clinical requirements. BASE and Reggie are freely available and released as open-source under the GNU General Public License and GNU Affero General Public License, respectively.\n\nConclusionUsing free open-source software together with BASE and a customized extension package, Reggie, we have implemented a system capable of managing large collections of quality controlled and curated material for use in research and development and tailored to meet requirements for clinical use. Featuring high degree of automation and interactivity the system allows for resource efficient laboratory procedures and short lead-times with demonstrated use of RNAseq data analyses in a clinical setting.

Bioinformatics

GenVisR: Genomic Visualizations in R

SummaryVisualizing and summarizing data from genomic studies continues to be a challenge. Here we introduce the GenVisR package to addresses this challenge by providing highly customizable, publication-quality graphics focused on cohort level genome analyses. GenVisR provides a rapid and easy-to-use suite of genomic visualization tools, while maintaining a high degree of flexibility by leveraging the abilities of ggplot2 and bioconductor.\n\nAvailability and ImplementationGenVisR is an R package available via bioconductor (https://bioconductor.org/packages/GenVisR) under GPLv3. Support is available via GitHub (https://github.com/griffithlab/GenVisR/issues) and the Bioconductor support website.\n\nContactogriffit@genome.wustl.edu, mgriffit@genome.wustl.edu

Bioinformatics

DiscoMark: Nuclear marker discovery from orthologous sequences using draft genome data

High-throughput sequencing has laid the foundation for fast and cost-effective development of phylogenetic markers. Here we present the program DO_SCPCAPISCOC_SCPCAPMO_SCPCAPARKC_SCPCAP, which streamlines the development of nuclear DNA (nDNA) markers from whole-genome (or whole-transcriptome) sequencing data, combining local alignment, alignment trimming, reference mapping and primer design based on multiple sequence alignments in order to design primer pairs from input orthologous sequences. In order to demonstrate the suitability of DO_SCPCAPISCOC_SCPCAPMO_SCPCAPARKC_SCPCAP we designed markers for two groups of species, one consisting of closely related species and one group of distantly related species. For the closely related members of the species complex of Cloeon dipterum s.l. (Insecta, Ephemeroptera), the program discovered a total of 78 markers. Among these, we selected eight markers for amplification and Sanger sequencing. The exon sequence alignments (2,526 base pairs (bp)) were used to reconstruct a well supported phylogeny and to infer clearly structured haplotype networks. For the distantly related species we designed primers for several families in the insect order Ephemeroptera, using available genomic data from four sequenced species. We developed primer pairs for 23 markers that are designed to amplify across several families. The DO_SCPCAPISCOC_SCPCAPMO_SCPCAPARKC_SCPCAP program will enhance the development of new nDNA markersby providing a streamlined, automated approach to perform genome-scale scans for phylogenetic markers. The program is written in Python, released under a public license (GNU GPL v2), and together with a manual and example data set available at: https://github.com/hdetering/discomark.

Bioinformatics