Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Mapping challenging mutations by whole-genome sequencing

Whole-genome sequencing provides a rapid and powerful method for identifying mutations on a global scale, and has spurred a renewed enthusiasm for classical genetic screens in model organisms. The most commonly characterized category of mutation consists of monogenic, recessive traits, due to their genetic tractability. Therefore, most of the mapping methods for mutation identification by whole-genome sequencing are directed toward alleles that fulfill those criteria (i.e., single-gene, homozygous variants). However, such approaches are not entirely suitable for the characterization of a variety of more challenging mutations, such as dominant and semi-dominant alleles or multigenic traits. Therefore, we have developed strategies for the identification of those classes of mutations, using polymorphism mapping in Caenorhabditis elegans as our model for validation. We also report an alternative approach for mutation identification from traditional recombinant crosses, and a solution to the technical challenge of sequencing sterile or terminally arrested strains where population size is limiting. The methods described herein extend the applicability of whole-genome sequencing to a broader spectrum of mutations, including classes that are difficult to map by traditional means.

Genetics

Tempo and mode of genome evolution in a 50,000-generation experiment

Adaptation depends on the rates, effects, and interactions of many mutations. We analyzed 264 genomes from 12 Escherichia coli populations to characterize their dynamics over 50,0 generations. The trajectories for genome evolution in populations that retained the ancestral mutation rate fit a model where most fixed mutations are beneficial, the fraction of beneficial mutations declines as fitness rises, and neutral mutations accumulate at a constant rate. We also compared these populations to lines evolved under a mutation-accumulation regime that minimizes selection. Nonsynonymous mutations, intergenic mutations, insertions, and deletions are overrepresented in the long-term populations, supporting the inference that most fixed mutations are favored by selection. These results illuminate the shifting balance of forces that govern genome evolution in populations adapting to a new environment.

Evolutionary Biology

Using Genome Wide Estimates of Heritability to Examine the Relevance of Gene-Environment Interplay

We use genome-wide data from the third generation respondents of the Framing-ham Heart Study to estimate heritability in body mass index using different quantities of the measured genotype. Heritability decreases rapidly when SNPs implicated by a genome-wide association study are removed but shows essentially no decline when SNPs implicated by a gene-environment interaction in a second genome-wide analysis are removed. This second result is highlighted by our additional finding that the SNPs which explain heritability amongst a subsample defined by higher educational attainment explain no heritability of the heritability in the lower education group, and vice-versa. Finally, we do find consistent heritability estimates when we compare family-based estimates versus those based on measured genotype.

Genetics

Phenotypic plasticity promotes balanced polymorphism in periodic environments by a genomic storage effect

Phenotypic plasticity is known to evolve in perturbed habitats, where it alleviates the deleterious effects of selection. But the effects of plasticity on levels of genetic polymorphism, an important precursor to adaptation in temporally varying environments, are unclear. Here we develop a haploid, two-locus population-genetic model to describe the interplay between a plasticity modifier locus and a target locus subject to periodically varying selection. We find that the interplay between these two loci can produce a \"genomic storage effect\" that promotes balanced polymorphism over a large range of parameters, in the absence of all other conditions known to maintain genetic variation. The genomic storage effect arises as recombination allows alleles at the two loci to escape more harmful genetic backgrounds and associate in haplotypes that persist until environmental conditions change. Using both Monte Carlo simulations and analytical approximations we quantify the strength of the genomic storage effect across a range of selection pressures, recombination rates, plasticity modifier effect sizes, and environmental periods.

Evolutionary Biology

Implementation of an Open Source Software solution for Laboratory Information Management and automated RNAseq data analysis in a large-scale Cancer Genomics initiative using BASE with extension package Reggie.

BackgroundLarge-scale cancer genomics initiatives and next-generation sequencing for transcriptome profiling allow for detailed molecular characterization of tumors, and provide opportunities for clinical tools to improve diagnosis, prognosis, and treatment decisions. Laboratory information, data management, and data sharing in large-scale genomics projects is a challenge. Aiming to introduce such technologies in a clinical setting offer additional challenges associated with requirements of short lead-times and specialized tracking of biomaterials, data, and analysis results.\n\nResultsUsing the free open-source BioArray Software Environment (BASE) and extension package Reggie we have implemented a laboratory information management system and an automated RNAseq data analysis pipeline that successfully manage a large regional cancer genomics initiative. The system manages enrolled cancer patients, tumor biopsies, extraction of nucleic acid, and whole transcriptome RNA-sequencing through to data analysis and quality control. The implementation offers integration of laboratory equipment and operating procedures, and information tracking in a module based fashion enabling efficient and flexible use of personnel resources. The system provides two-factor authentication and transaction control and seamless integration of freely available software for RNAseq analysis such as Tophat, Cufflinks, and Picard. As of February 2016 more than 8000 patients and over 6000 tumor biopsies have been successfully processed. Lead-time from biopsy arrival to summarized reports based on RNAseq data is less than 5 days, in line with regional clinical requirements. BASE and Reggie are freely available and released as open-source under the GNU General Public License and GNU Affero General Public License, respectively.\n\nConclusionUsing free open-source software together with BASE and a customized extension package, Reggie, we have implemented a system capable of managing large collections of quality controlled and curated material for use in research and development and tailored to meet requirements for clinical use. Featuring high degree of automation and interactivity the system allows for resource efficient laboratory procedures and short lead-times with demonstrated use of RNAseq data analyses in a clinical setting.

Bioinformatics

GenVisR: Genomic Visualizations in R

SummaryVisualizing and summarizing data from genomic studies continues to be a challenge. Here we introduce the GenVisR package to addresses this challenge by providing highly customizable, publication-quality graphics focused on cohort level genome analyses. GenVisR provides a rapid and easy-to-use suite of genomic visualization tools, while maintaining a high degree of flexibility by leveraging the abilities of ggplot2 and bioconductor.\n\nAvailability and ImplementationGenVisR is an R package available via bioconductor (https://bioconductor.org/packages/GenVisR) under GPLv3. Support is available via GitHub (https://github.com/griffithlab/GenVisR/issues) and the Bioconductor support website.\n\nContactogriffit@genome.wustl.edu, mgriffit@genome.wustl.edu

Bioinformatics

DiscoMark: Nuclear marker discovery from orthologous sequences using draft genome data

High-throughput sequencing has laid the foundation for fast and cost-effective development of phylogenetic markers. Here we present the program DO_SCPCAPISCOC_SCPCAPMO_SCPCAPARKC_SCPCAP, which streamlines the development of nuclear DNA (nDNA) markers from whole-genome (or whole-transcriptome) sequencing data, combining local alignment, alignment trimming, reference mapping and primer design based on multiple sequence alignments in order to design primer pairs from input orthologous sequences. In order to demonstrate the suitability of DO_SCPCAPISCOC_SCPCAPMO_SCPCAPARKC_SCPCAP we designed markers for two groups of species, one consisting of closely related species and one group of distantly related species. For the closely related members of the species complex of Cloeon dipterum s.l. (Insecta, Ephemeroptera), the program discovered a total of 78 markers. Among these, we selected eight markers for amplification and Sanger sequencing. The exon sequence alignments (2,526 base pairs (bp)) were used to reconstruct a well supported phylogeny and to infer clearly structured haplotype networks. For the distantly related species we designed primers for several families in the insect order Ephemeroptera, using available genomic data from four sequenced species. We developed primer pairs for 23 markers that are designed to amplify across several families. The DO_SCPCAPISCOC_SCPCAPMO_SCPCAPARKC_SCPCAP program will enhance the development of new nDNA markersby providing a streamlined, automated approach to perform genome-scale scans for phylogenetic markers. The program is written in Python, released under a public license (GNU GPL v2), and together with a manual and example data set available at: https://github.com/hdetering/discomark.

Bioinformatics

Genome-wide generalized additive models

MotivationChromatin immunoprecipitation followed by deep sequencing (ChIP-Seq) is a widely used approach to study protein-DNA interactions. Often, the quantities of interest are the differential occupancies relative to controls, between genetic backgrounds, treatments, or combinations thereof. Current methods for differential occupancy of ChIP-seq data rely however on binning or sliding window techniques, for which the choice of the window and bin sizes are subjective.\n\nResultsHere, we present GenoGAM (Genome-wide Generalized Additive Model), which brings the well-established and flexible generalized additive models framework to genomic applications using a data parallelism strategy. We model ChIP-Seq read count frequencies as products of smooth functions along chromosomes. Smoothing parameters are objectively estimated from the data by cross-validation, eliminating ad-hoc binning and windowing needed by current approaches. GenoGAM provides base-level and region-level significance testing for full factorial designs. Application to a ChIP-Seq dataset in yeast showed increased sensitivity over existing differential occupancy methods while controlling for type I error rate. By analyzing a set of DNA methylation data and illustrating an extension to a peak caller, we further demonstrate the potential of GenoGAM as a generic statistical modeling tool for genome-wide assays.\n\nAvailabilitySoftware is available from Bioconductor: https://www.bioconductor.org/packages/release/bioc/html/GenoGAM.html\n\nContactgagneur@in.tum.de\n\nSupplementary informationSupplementary information is available at Bioinformatics online.

Bioinformatics

Do aye-ayes echolocate? Studying convergent genomic evolution in a primate auditory specialist

Several taxonomically distinct mammalian groups - certain microbats and cetaceans (e.g. dolphins) - share both morphological adaptations related to echolocation behavior and strong signatures of convergent evolution at the amino acid level across seven genes related to auditory processing. Aye-ayes (Daubentonia madagascariensis) are nocturnal lemurs with a derived auditory processing system. Aye-ayes tap rapidly along the surfaces of dead trees, listening to reverberations to identify the mines of wood-boring insect larvae; this behavior has been hypothesized to functionally mimic echolocation. Here we investigated whether there are signals of genomic convergence between aye-ayes and known mammalian echolocators. We developed a computational pipeline (BEAT: Basic Exon Assembly Tool) that produces consensus sequences for regions of interest from shotgun genomic sequencing data for non-model organisms without requiring de novo genome assembly. We reconstructed complete coding region sequences for the seven convergent echolocating bat-dolphin genes for aye-ayes and another lemur. Sequences were compared in a phylogenetic framework to those of bat and dolphin echolocators and appropriate non-echolocating outgroups. Our analysis reaffirms the existence of amino acid convergence at these loci among echolocating bats and dolphins; we also detected unexpected signals of convergence between echolocating bats and both mice and elephants. However, we observed no significant signal of amino acid convergence between aye-ayes and echolocating bats and dolphins; our results thus suggest that aye-aye tap-foraging auditory adaptations represent distinct evolutionary innovations. These results are also consistent with a developing consensus that convergent behavioral ecology is not necessarily a reliable guide to convergent molecular evolution.

Evolutionary Biology

plasmidSPAdes: Assembling Plasmids from Whole Genome Sequencing Data

MotivationPlasmids are stably maintained extra-chromosomal genetic elements that replicate independently from the host cells chromosomes. Although plasmids harbor biomedically important genes, (such as genes involved in virulence and antibiotics resistance), there is a shortage of specialized software tools for extracting and assembling plasmid data from whole genome sequencing projects.\n\nResultsWe present the plasmidSPAdes algorithm and software tool for assembling plasmids from whole genome sequencing data and benchmark its performance on a diverse set of bacterial genomes.\n\nAvailability and implementationO_SCPCAPPLASMIDC_SCPCAPSPAO_SCPCAPDESC_SCPCAP is publicly available at http://spades.bioinf.spbau.ru/plasmidSPAdes/\n\nContactd.antipov@spbu.ru

Bioinformatics

Recombineering in C. elegans: genome editing using in vivo assembly of linear DNAs

Recombineering, the use of endogenous homologous recombination systems to recombine DNA in vivo, is a commonly used technique for genome editing in microbes. Recombineering has not yet been developed for animals, where non-homology-based mechanisms have been thought to dominate DNA repair. Here, we demonstrate that homology-dependent repair (HDR) is robust in C. elegans using linear templates with short homologies (~35 bases). Templates with homology to only one side of a double-strand break initiate repair efficiently, and short overlaps between templates support template switching. We demonstrate the use of single-stranded, bridging oligonucleotides (ssODNs) to target PCR fragments precisely to DSBs induced by CRISPR/Cas9 on chromosomes. Based on these findings, we develop recombineering strategies for genome editing that expand the utility of ssODNs and eliminate in vitro cloning steps for template construction. We apply these methods to the generation of GFP knock-in alleles and gene replacements without co-integrated markers. We conclude that, like microbes, metazoans possess robust homology-dependent repair mechanisms that can be harnessed for recombineering and genome editing.

Bioengineering

Homeostatic responses regulate selfish mitochondrial genome dynamics in C. elegans

Selfish genetic elements have profound biological and evolutionary consequences. Mutant mitochondrial genomes (mtDNA) can be viewed as selfish genetic elements that persist in a state of heteroplasmy despite having potentially deleterious consequences to the organism. We sought to investigate mechanisms that allow selfish mtDNA to achieve and sustain high levels. Here, we establish a large 3.1kb deletion bearing mtDNA variant uaDf5 as a bona fide selfish genome in the nematode Caenorhabditis elegans. Next, using droplet digital PCR to quantify mtDNA copy number, we show that uaDf5 mutant mtDNA replicates in addition to, not at the expense of, wildtype mtDNA. These data suggest existence of homeostatic copy number control for wildtype mtDNA that is exploited by uaDf5 to hitchhike to high frequency. We also observe activation of the mitochondrial unfolded protein response (UPRmt) in animals with uaDf5. Loss of UPRmt results in a decrease in uaDf5 frequency whereas constitutive activation of UPRmt increases uaDf5 levels. These data suggest that UPRmt allows uaDf5 levels to increase. Interestingly, the decreased uaDf5 levels in absence of UPRmt recover in parkin mutants lacking mitophagy, suggesting that UPRmt protects uaDf5 from mitophagy. We propose that cells activate two homeostatic responses, mtDNA copy number control and UPRmt, in uaDf5 heteroplasmic animals. Inadvertently, these homeostatic responses allow uaDf5 levels to be higher than they would be otherwise. In conclusion, our data suggest that homeostatic stress response mechanisms play an important role in regulating selfish mitochondrial genome dynamics.

Genetics

Advances in the integration of transcriptional regulatory information into genome-scale metabolic models

A major goal of systems biology is to build predictive computational models of cellular metabolism. Availability of complete genome sequences and wealth of legacy biochemical information has led to the reconstruction of genome-scale metabolic networks in the last 15 years for several organisms across the three domains of life. Due to paucity of information on kinetic parameters associated with metabolic reactions, the constraint-based modelling approach, flux balance analysis (FBA), has proved to be a vital alternative to investigate the capabilities of reconstructed metabolic networks. In parallel, advent of high-throughput technologies has led to the generation of massive amounts of omics data on transcriptional regulation comprising mRNA transcript levels and genome-wide binding profile of transcriptional regulators. A frontier area in metabolic systems biology has been the development of methods to integrate the available transcriptional regulatory information into constraint-based models of reconstructed metabolic networks in order to increase the predictive capabilities of computational models and understand the regulation of cellular metabolism. Here, we review the existing methods to integrate transcriptional regulatory information into constraint-based models of metabolic networks.

Systems Biology

A better design for stratified medicine based on genomic prediction

Genomic prediction shows promise for personalised medicine in which diagnosis and treatment are tailored to individuals based on their genetic profiles. Genomic prediction is arguably the greatest need for complex diseases and disorders for which both genetic and non-genetic factors contribute to risk. However, we have no adequate insight of the accuracy of such predictions, and how accuracy may vary between individuals or between populations. In this study, we present a theoretical framework to demonstrate that prediction accuracy can be maximised by targeting more informative individuals in a discovery set with closer relationships with the subjects, making prediction more similar to those in populations with small effective size (Ne). Increase of prediction accuracy from closer relationships is achieved under an additive model and does not rely on any interaction effects (gene x gene, gene x environment or gene x family). Using theory, simulations and real data analyses, we show that the predictive accuracy or the area under the receiver operating characteristic curve (AUC) increased exponentially with decreasing Ne. For example, with a set of realistic parameters (the sample size of discovery set N=3000 and heritability h2=0.5), AUC value approached to 0.9 (Ne=100) from 0.6 (Ne=10000), and the top percentile of the estimated genetic profile scores had 23 times higher proportion of cases than the general population (with Ne=100), which increased from 2 times higher proportion of cases (with Ne=10000). This suggests that different interventions in the top percentile risk groups maybe justified (i.e. stratified medicine). In conclusion, it is argued that there is considerable room to increase prediction accuracy for polygenic traits by using an efficient design of a smaller Ne (e.g. a design consisting of closer relationships) so that genomic prediction can be more beneficial in clinical applications in the near future.

Genetics

Phased Diploid Genome Assembly with Single Molecule Real-Time Sequencing

While genome assembly projects have been successful in a number of haploid or inbred species, one of the current main challenges is assembling non-inbred or rearranged heterozygous genomes. To address this critical need, we introduce the open-source FALCON and FALCON-Unzip algorithms (https://github.com/PacificBiosciences/FALCON/) to assemble Single Molecule Real-Time (SMRT(R)) Sequencing data into highly accurate, contiguous, and correctly phased diploid genomes. We demonstrate the quality of this approach by assembling new reference sequences for three heterozygous samples, including an F1 hybrid of the model species Arabidopsis thaliana, the widely cultivated V. vinifera cv. Cabernet Sauvignon, and the coral fungus Clavicorona pyxidata that have challenged short-read assembly approaches. The FALCON-based assemblies were substantially more contiguous and complete than alternate short or long-read approaches. The phased diploid assembly enabled the study of haplotype structures and heterozygosities between the homologous chromosomes, including identifying widespread heterozygous structural variations within the coding sequences.

Bioinformatics

Estimating the functional impact of INDELs in transcription factor binding sites: a genome-wide landscape

BackgroundVariants in transcription factor binding sites (TFBSs) may have important regulatory effects, as they have the potential to alter transcription factor (TF) binding affinities and thereby affecting gene expression. With recent advances in sequencing technologies the number of variants identified in TFBSs has increased, hence understanding their role is of significant interest when interpreting next generation sequencing data. Current methods have two major limitations: they are limited to predicting the functional impact of single nucleotide variants (SNVs) and often rely on additional experimental data, laborious and expensive to acquire. We propose a purely bioinformatic method that addresses these two limitations while providing comparable results.\n\nResultsOur method uses position weight matrices and a sliding window approach, in order to account for the sequence context of variants, and scores the consequences of both SNVs and INDELs in TFBSs. We tested the accuracy of our method in two different ways. Firstly, we compared it to a recent method based on DNase I hypersensitive sites sequencing (DHS-seq) data designed to predict the effects of SNVs: we found a significant correlation of our score both with their DHS-seq data and their prediction model. Secondly, we called INDELs on publicly available DHS-seq data from ENCODE, and found our score to represent well the experimental data. We concluded that our method is reliable and we used it to describe the landscape of variation in TFBSs in the human genome, by scoring all variants in the 1000 Genomes Project Phase 3. Surprisingly, we found that most insertions have neutral effects on binding sites, while deletions, as expected, were found to have the most severe TFBS-scores. We identified four categories of variants based on their TFBS-scores and tested them for enrichment of variants classified as pathogenic, benign and protective in ClinVar: we found that the variants with the most negative TFBS-scores have the most significant enrichment for pathogenic variants.\n\nConclusionsOur method addresses key shortcomings of currently available bioinformatic tools in predicting the effects of INDELs in TFBSs, and provides an unprecedented window into the genome-wide landscape of INDELs, their predicted influences on TF binding, and potential relevance for human diseases. We thus offer an additional tool to help prioritising non-coding variants in sequencing studies.

Bioinformatics

Post-selection Inference Following Aggregate Level Hypothesis Testing in Large Scale Genomic Data

In many genomic applications, hypotheses tests are performed by aggregating test-statistics across units within naturally defined classes for powerful identification of signals. Following class-level testing, it is naturally of interest to identify the lower level units which contain true signals. Testing the individual units within a class without taking into account the fact that the class was selected using an aggregate-level test-statistic, will produce biased inference. We develop a hypothesis testing framework that guarantees control for false positive rates conditional on the fact that the class was selected. Specifically, we develop procedures for calculating unit level p-values that allows rejection of null hypotheses controlling for two types of conditional error rates, one relating to family wise rate and the other relating to false discovery rate. We use simulation studies to illustrate validity and power of the proposed procedure in comparison to several possible alternatives. We illustrate the power of the method in a natural application involving whole-genome expression quantitative trait loci (eQTL) analysis across 17 tissue types using data from The Cancer Genome Atlas (TCGA) Project.

Bioinformatics

Deciphering the wisent demographic and adaptive histories from individual whole-genome sequences

As the largest European herbivore, the wisent (Bison bonasus) is emblematic of the continent wildlife but has unclear origins. Here, we infer its demographic and adaptive histories from two individual whole genome sequences via a detailed comparative analysis with bovine genomes. We estimate that the wisent and bovine species diverged from 1.7x106 to 850,000 YBP through a speciation process involving an extended period of limited gene flow. Our data further support the occurrence of more recent secondary contacts, posterior to the Bos taurus and Bos indicus divergence (ca. 150,000 YBP), between the wisent and (European) taurine cattle lineages. Although the wisent and bovine population sizes experienced a similar sharp decline since the Last Glacial Maximum, we find that the wisent demography remained more fluctuating during the Pleistocene. This is in agreement with a scenario in which wisents responded to successive glaciations by habitat fragmentation rather than southward and eastward migration as for the bovine ancestors.\n\nWe finally detect 423 genes under positive selection between the wisent and bovine lineages, which shed a new light on the genome response to different living conditions (temperature, available food resource and pathogen exposure) and on the key gene functions altered by the domestication process.

Evolutionary Biology