Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

Genome-wide generalized additive models

MotivationChromatin immunoprecipitation followed by deep sequencing (ChIP-Seq) is a widely used approach to study protein-DNA interactions. Often, the quantities of interest are the differential occupancies relative to controls, between genetic backgrounds, treatments, or combinations thereof. Current methods for differential occupancy of ChIP-seq data rely however on binning or sliding window techniques, for which the choice of the window and bin sizes are subjective.\n\nResultsHere, we present GenoGAM (Genome-wide Generalized Additive Model), which brings the well-established and flexible generalized additive models framework to genomic applications using a data parallelism strategy. We model ChIP-Seq read count frequencies as products of smooth functions along chromosomes. Smoothing parameters are objectively estimated from the data by cross-validation, eliminating ad-hoc binning and windowing needed by current approaches. GenoGAM provides base-level and region-level significance testing for full factorial designs. Application to a ChIP-Seq dataset in yeast showed increased sensitivity over existing differential occupancy methods while controlling for type I error rate. By analyzing a set of DNA methylation data and illustrating an extension to a peak caller, we further demonstrate the potential of GenoGAM as a generic statistical modeling tool for genome-wide assays.\n\nAvailabilitySoftware is available from Bioconductor: https://www.bioconductor.org/packages/release/bioc/html/GenoGAM.html\n\nContactgagneur@in.tum.de\n\nSupplementary informationSupplementary information is available at Bioinformatics online.

Bioinformatics

Do aye-ayes echolocate? Studying convergent genomic evolution in a primate auditory specialist

Several taxonomically distinct mammalian groups - certain microbats and cetaceans (e.g. dolphins) - share both morphological adaptations related to echolocation behavior and strong signatures of convergent evolution at the amino acid level across seven genes related to auditory processing. Aye-ayes (Daubentonia madagascariensis) are nocturnal lemurs with a derived auditory processing system. Aye-ayes tap rapidly along the surfaces of dead trees, listening to reverberations to identify the mines of wood-boring insect larvae; this behavior has been hypothesized to functionally mimic echolocation. Here we investigated whether there are signals of genomic convergence between aye-ayes and known mammalian echolocators. We developed a computational pipeline (BEAT: Basic Exon Assembly Tool) that produces consensus sequences for regions of interest from shotgun genomic sequencing data for non-model organisms without requiring de novo genome assembly. We reconstructed complete coding region sequences for the seven convergent echolocating bat-dolphin genes for aye-ayes and another lemur. Sequences were compared in a phylogenetic framework to those of bat and dolphin echolocators and appropriate non-echolocating outgroups. Our analysis reaffirms the existence of amino acid convergence at these loci among echolocating bats and dolphins; we also detected unexpected signals of convergence between echolocating bats and both mice and elephants. However, we observed no significant signal of amino acid convergence between aye-ayes and echolocating bats and dolphins; our results thus suggest that aye-aye tap-foraging auditory adaptations represent distinct evolutionary innovations. These results are also consistent with a developing consensus that convergent behavioral ecology is not necessarily a reliable guide to convergent molecular evolution.

Evolutionary Biology

plasmidSPAdes: Assembling Plasmids from Whole Genome Sequencing Data

MotivationPlasmids are stably maintained extra-chromosomal genetic elements that replicate independently from the host cells chromosomes. Although plasmids harbor biomedically important genes, (such as genes involved in virulence and antibiotics resistance), there is a shortage of specialized software tools for extracting and assembling plasmid data from whole genome sequencing projects.\n\nResultsWe present the plasmidSPAdes algorithm and software tool for assembling plasmids from whole genome sequencing data and benchmark its performance on a diverse set of bacterial genomes.\n\nAvailability and implementationO_SCPCAPPLASMIDC_SCPCAPSPAO_SCPCAPDESC_SCPCAP is publicly available at http://spades.bioinf.spbau.ru/plasmidSPAdes/\n\nContactd.antipov@spbu.ru

Bioinformatics

Recombineering in C. elegans: genome editing using in vivo assembly of linear DNAs

Recombineering, the use of endogenous homologous recombination systems to recombine DNA in vivo, is a commonly used technique for genome editing in microbes. Recombineering has not yet been developed for animals, where non-homology-based mechanisms have been thought to dominate DNA repair. Here, we demonstrate that homology-dependent repair (HDR) is robust in C. elegans using linear templates with short homologies (~35 bases). Templates with homology to only one side of a double-strand break initiate repair efficiently, and short overlaps between templates support template switching. We demonstrate the use of single-stranded, bridging oligonucleotides (ssODNs) to target PCR fragments precisely to DSBs induced by CRISPR/Cas9 on chromosomes. Based on these findings, we develop recombineering strategies for genome editing that expand the utility of ssODNs and eliminate in vitro cloning steps for template construction. We apply these methods to the generation of GFP knock-in alleles and gene replacements without co-integrated markers. We conclude that, like microbes, metazoans possess robust homology-dependent repair mechanisms that can be harnessed for recombineering and genome editing.

Bioengineering

Homeostatic responses regulate selfish mitochondrial genome dynamics in C. elegans

Selfish genetic elements have profound biological and evolutionary consequences. Mutant mitochondrial genomes (mtDNA) can be viewed as selfish genetic elements that persist in a state of heteroplasmy despite having potentially deleterious consequences to the organism. We sought to investigate mechanisms that allow selfish mtDNA to achieve and sustain high levels. Here, we establish a large 3.1kb deletion bearing mtDNA variant uaDf5 as a bona fide selfish genome in the nematode Caenorhabditis elegans. Next, using droplet digital PCR to quantify mtDNA copy number, we show that uaDf5 mutant mtDNA replicates in addition to, not at the expense of, wildtype mtDNA. These data suggest existence of homeostatic copy number control for wildtype mtDNA that is exploited by uaDf5 to hitchhike to high frequency. We also observe activation of the mitochondrial unfolded protein response (UPRmt) in animals with uaDf5. Loss of UPRmt results in a decrease in uaDf5 frequency whereas constitutive activation of UPRmt increases uaDf5 levels. These data suggest that UPRmt allows uaDf5 levels to increase. Interestingly, the decreased uaDf5 levels in absence of UPRmt recover in parkin mutants lacking mitophagy, suggesting that UPRmt protects uaDf5 from mitophagy. We propose that cells activate two homeostatic responses, mtDNA copy number control and UPRmt, in uaDf5 heteroplasmic animals. Inadvertently, these homeostatic responses allow uaDf5 levels to be higher than they would be otherwise. In conclusion, our data suggest that homeostatic stress response mechanisms play an important role in regulating selfish mitochondrial genome dynamics.

Genetics

Advances in the integration of transcriptional regulatory information into genome-scale metabolic models

A major goal of systems biology is to build predictive computational models of cellular metabolism. Availability of complete genome sequences and wealth of legacy biochemical information has led to the reconstruction of genome-scale metabolic networks in the last 15 years for several organisms across the three domains of life. Due to paucity of information on kinetic parameters associated with metabolic reactions, the constraint-based modelling approach, flux balance analysis (FBA), has proved to be a vital alternative to investigate the capabilities of reconstructed metabolic networks. In parallel, advent of high-throughput technologies has led to the generation of massive amounts of omics data on transcriptional regulation comprising mRNA transcript levels and genome-wide binding profile of transcriptional regulators. A frontier area in metabolic systems biology has been the development of methods to integrate the available transcriptional regulatory information into constraint-based models of reconstructed metabolic networks in order to increase the predictive capabilities of computational models and understand the regulation of cellular metabolism. Here, we review the existing methods to integrate transcriptional regulatory information into constraint-based models of metabolic networks.

Systems Biology

A better design for stratified medicine based on genomic prediction

Genomic prediction shows promise for personalised medicine in which diagnosis and treatment are tailored to individuals based on their genetic profiles. Genomic prediction is arguably the greatest need for complex diseases and disorders for which both genetic and non-genetic factors contribute to risk. However, we have no adequate insight of the accuracy of such predictions, and how accuracy may vary between individuals or between populations. In this study, we present a theoretical framework to demonstrate that prediction accuracy can be maximised by targeting more informative individuals in a discovery set with closer relationships with the subjects, making prediction more similar to those in populations with small effective size (Ne). Increase of prediction accuracy from closer relationships is achieved under an additive model and does not rely on any interaction effects (gene x gene, gene x environment or gene x family). Using theory, simulations and real data analyses, we show that the predictive accuracy or the area under the receiver operating characteristic curve (AUC) increased exponentially with decreasing Ne. For example, with a set of realistic parameters (the sample size of discovery set N=3000 and heritability h2=0.5), AUC value approached to 0.9 (Ne=100) from 0.6 (Ne=10000), and the top percentile of the estimated genetic profile scores had 23 times higher proportion of cases than the general population (with Ne=100), which increased from 2 times higher proportion of cases (with Ne=10000). This suggests that different interventions in the top percentile risk groups maybe justified (i.e. stratified medicine). In conclusion, it is argued that there is considerable room to increase prediction accuracy for polygenic traits by using an efficient design of a smaller Ne (e.g. a design consisting of closer relationships) so that genomic prediction can be more beneficial in clinical applications in the near future.

Genetics

Phased Diploid Genome Assembly with Single Molecule Real-Time Sequencing

While genome assembly projects have been successful in a number of haploid or inbred species, one of the current main challenges is assembling non-inbred or rearranged heterozygous genomes. To address this critical need, we introduce the open-source FALCON and FALCON-Unzip algorithms (https://github.com/PacificBiosciences/FALCON/) to assemble Single Molecule Real-Time (SMRT(R)) Sequencing data into highly accurate, contiguous, and correctly phased diploid genomes. We demonstrate the quality of this approach by assembling new reference sequences for three heterozygous samples, including an F1 hybrid of the model species Arabidopsis thaliana, the widely cultivated V. vinifera cv. Cabernet Sauvignon, and the coral fungus Clavicorona pyxidata that have challenged short-read assembly approaches. The FALCON-based assemblies were substantially more contiguous and complete than alternate short or long-read approaches. The phased diploid assembly enabled the study of haplotype structures and heterozygosities between the homologous chromosomes, including identifying widespread heterozygous structural variations within the coding sequences.

Bioinformatics

Estimating the functional impact of INDELs in transcription factor binding sites: a genome-wide landscape

BackgroundVariants in transcription factor binding sites (TFBSs) may have important regulatory effects, as they have the potential to alter transcription factor (TF) binding affinities and thereby affecting gene expression. With recent advances in sequencing technologies the number of variants identified in TFBSs has increased, hence understanding their role is of significant interest when interpreting next generation sequencing data. Current methods have two major limitations: they are limited to predicting the functional impact of single nucleotide variants (SNVs) and often rely on additional experimental data, laborious and expensive to acquire. We propose a purely bioinformatic method that addresses these two limitations while providing comparable results.\n\nResultsOur method uses position weight matrices and a sliding window approach, in order to account for the sequence context of variants, and scores the consequences of both SNVs and INDELs in TFBSs. We tested the accuracy of our method in two different ways. Firstly, we compared it to a recent method based on DNase I hypersensitive sites sequencing (DHS-seq) data designed to predict the effects of SNVs: we found a significant correlation of our score both with their DHS-seq data and their prediction model. Secondly, we called INDELs on publicly available DHS-seq data from ENCODE, and found our score to represent well the experimental data. We concluded that our method is reliable and we used it to describe the landscape of variation in TFBSs in the human genome, by scoring all variants in the 1000 Genomes Project Phase 3. Surprisingly, we found that most insertions have neutral effects on binding sites, while deletions, as expected, were found to have the most severe TFBS-scores. We identified four categories of variants based on their TFBS-scores and tested them for enrichment of variants classified as pathogenic, benign and protective in ClinVar: we found that the variants with the most negative TFBS-scores have the most significant enrichment for pathogenic variants.\n\nConclusionsOur method addresses key shortcomings of currently available bioinformatic tools in predicting the effects of INDELs in TFBSs, and provides an unprecedented window into the genome-wide landscape of INDELs, their predicted influences on TF binding, and potential relevance for human diseases. We thus offer an additional tool to help prioritising non-coding variants in sequencing studies.

Bioinformatics

Post-selection Inference Following Aggregate Level Hypothesis Testing in Large Scale Genomic Data

In many genomic applications, hypotheses tests are performed by aggregating test-statistics across units within naturally defined classes for powerful identification of signals. Following class-level testing, it is naturally of interest to identify the lower level units which contain true signals. Testing the individual units within a class without taking into account the fact that the class was selected using an aggregate-level test-statistic, will produce biased inference. We develop a hypothesis testing framework that guarantees control for false positive rates conditional on the fact that the class was selected. Specifically, we develop procedures for calculating unit level p-values that allows rejection of null hypotheses controlling for two types of conditional error rates, one relating to family wise rate and the other relating to false discovery rate. We use simulation studies to illustrate validity and power of the proposed procedure in comparison to several possible alternatives. We illustrate the power of the method in a natural application involving whole-genome expression quantitative trait loci (eQTL) analysis across 17 tissue types using data from The Cancer Genome Atlas (TCGA) Project.

Bioinformatics

Deciphering the wisent demographic and adaptive histories from individual whole-genome sequences

As the largest European herbivore, the wisent (Bison bonasus) is emblematic of the continent wildlife but has unclear origins. Here, we infer its demographic and adaptive histories from two individual whole genome sequences via a detailed comparative analysis with bovine genomes. We estimate that the wisent and bovine species diverged from 1.7x106 to 850,000 YBP through a speciation process involving an extended period of limited gene flow. Our data further support the occurrence of more recent secondary contacts, posterior to the Bos taurus and Bos indicus divergence (ca. 150,000 YBP), between the wisent and (European) taurine cattle lineages. Although the wisent and bovine population sizes experienced a similar sharp decline since the Last Glacial Maximum, we find that the wisent demography remained more fluctuating during the Pleistocene. This is in agreement with a scenario in which wisents responded to successive glaciations by habitat fragmentation rather than southward and eastward migration as for the bovine ancestors.\n\nWe finally detect 423 genes under positive selection between the wisent and bovine lineages, which shed a new light on the genome response to different living conditions (temperature, available food resource and pathogen exposure) and on the key gene functions altered by the domestication process.

Evolutionary Biology

A genomic view of the peopling of the Americas

Whole-genome studies have documented that most Native American ancestry stems from a single population that diversified within the continent more than twelve thousand years ago. However, this shared ancestry hides a more complex history whereby at least four distinct streams of Eurasian migration have contributed to present-day and prehistoric Native American populations. Whole genome studies enhanced by technological breakthroughs in ancient DNA now provide evidence of a sequence of events involving initial migration from a structured Northeast Asian source population, followed by a divergence into northern and southern Native American lineages. During the Holocene, new migrations from Asia introduced the Saqqaq/Dorset Paleoeskimo population to the North American Arctic ~4,500 years ago, ancestry that is potentially connected with ancestry found in Athabaskan-speakers today. This was then followed by a major new population turnover in the high Arctic involving Thule-related peoples who are the ancestors of present-day Inuit. We highlight several open questions that could be addressed through future genomic research.

Evolutionary Biology

A natural encoding of genetic variation in a Burrows-Wheeler Transform to enable mapping and genome inference

We show how positional markers can be used to encode genetic variation within aBurrows-Wheeler Transform (BWT), and use this to construct a generalisation ofthe traditional \"reference genome\", incorporating known variation within aspecies. Our goal is to support the inference of the closest mosaic of previouslyknown sequences to the genome(s) under analysis.\n\nOur scheme results in an increased alphabet size, and by using a wavelet tree encoding of the BWT we reduce the performance impact on rank operations. We give a specialised form of the backward search that allows variation-aware exact matching. We implement this, and demonstrate the cost of constructing an index of the whole human genome with 8 million genetic variants is 25GB of RAM. We also show that inferring a closer reference can close large kilobase-scale coverage gaps in P. falciparum.

Bioinformatics

Protocol: Genome-scale CRISPR-Cas9 Knockout and Transcriptional Activation Screening

Forward genetic screens are powerful tools for the unbiased discovery and functional characterization of specific genetic elements associated with a phenotype of interest. Recently, the RNA-guided endonuclease Cas9 from the microbial immune system CRISPR (clustered regularly interspaced short palindromic repeats) has been adapted for genome-scale screening by combining Cas9 with guide RNA libraries. Here we describe a protocol for genome-scale knockout and transcriptional activation screening using the CRISPR-Cas9 system. Custom-or ready-made guide RNA libraries are constructed and packaged into lentivirus for delivery into cells for screening. As each screen is unique, we provide guidelines for determining screening parameters and maintaining sufficient coverage. To validate candidate genes identified from the screen, we further describe strategies for confirming the screening phenotype as well as genetic perturbation through analysis of indel rate and transcriptional activation. Beginning with library design, a genome-scale screen can be completed in 6-10 weeks followed by 3-4 weeks of validation.

Molecular Biology

Comparative analysis highlights variable genome content of wheat rusts and divergence of the mating loci.

Three members of the Puccini genus, P. triticina (Pt), P. striiformis f.sp. tritici(Pst), and P. graminis f.sp. tritici (Pgt), cause the most common and often most significant foliar diseases of wheat. While similar in biology and life cycle, each species is uniquely adapted and specialized. The genomes of Pt and Pst were sequenced and compared to that of Pgt to identify common and distinguishing gene content, to determine gene variation among wheat rust pathogens, other rust fungi and basidiomycetes, and to identify genes of significance for infection. Pt had the largest genome of the three, estimated at 135 Mb with expansion due to mobile elements and repeats encompassing 50.9% of contig bases; by comparison repeats occupy 31.5% for Pst and 36.5% for Pgt. We find all three genomes are highly heterozygous, with Pst (5.97 SNPs/kb) nearly twice the level detected in Pt (2.57 SNPs/kb) and that previously reported for Pgt. Of 1,358 predicted effectors in Pt, 784 were found expressed across diverse life cycle stages including the sexual stage. Comparison to related fungi highlighted the expansion of gene families involved in transcriptional regulation and nucleotide binding, protein modification, and carbohydrate enzyme degradation. Two allelic homeodomain, HD1 and HD2, pairs and three pheromone receptor (STE3) mating-type genes were identified in each dikaryotic Puccinia species. The HD proteins were active in a heterologous Ustilago maydis mating assay and host induced gene silencing of the HD and STE3 alleles reduced wheat host infection.

Microbiology

Cross-species genome-wide identification of evolutionary conserved microProteins

MicroProteins are small single domain proteins that act by engaging their targets into non-productive protein complexes. In order to identify novel microProteins in any sequenced genome of interest, we have developed miPFinder, a program that identifies and classifies potential microProteins. In the past years, several microProteins have been discovered in plants where they are mainly involved in the regulation of development. The miPFinder algorithm identifies all up to date known plant microProteins and extends the microProtein concept to other protein families. Here, we reveal potential microProtein candidates in several plant and animal reference genomes. A large number of these microProteins are species-specific while others evolved early and are evolutionary highly conserved. Most known microProtein genes originated from large ancestral genes by gene duplication, mutation and subsequent degradation. Gene ontology analysis shows that putative microProtein ancestors are often located in the nucleus, and involved in DNA binding and formation of protein complexes. Additionally, microProtein candidates act in plant transcriptional regulation, signal transduction and anatomical structure development. MiPFinder is freely available to find microProteins in any genome and will aid in the identification of novel microProteins in plants and animals

Bioinformatics

Modelling the transcription factor DNA-binding affinity using genome-wide ChIP-based data

Understanding protein-DNA binding affinity is still a mystery for many transcription factors (TFs). Although several approaches have been proposed in the literature to model the DNA-binding specificity of TFs, they still have some limitations. Most of the methods require a cut-off threshold in order to classify a K-mer as a binding site (BS) and finding such a threshold is usually done by handcraft rather than a science. Some other approaches use a prior knowledge on the biological context of regulatory elements in the genome along with machine learning algorithms to build classifier models for TFBSs. Noticeably, these methods deliberately select the training and testing datasets so that they are very separable. Hence, the current methods do not actually capture the TF-DNA binding relationship. In this paper, we present a threshold-free framework based on a novel ensemble learning algorithm in order to locate TFBSs in DNA sequences. Our proposed approach creates TF-specific classifier models using genome-wide DNA-binding experiments and a prior biological knowledge on DNA sequences and TF binding preferences. Systematic background filtering algorithms are utilized to remove non-functional K-mers from training and testing datasets. To reduce the complexity of classifier models, a fast feature selection algorithm is employed. Finally, the created classifier models are used to scan new DNA sequences and identify potential binding sites. The analysis results show that our proposed approach is able to identify novel binding sites in the Saccharomyces cerevisiae genome.\n\nContactmonther.alhamdoosh@unimelb.edu.au, dh.wang@latrobe.edu.au\n\nAvailabilityhttp://homepage.cs.latrobe.edu.au/dwang/DNNESCANweb

Bioinformatics

KAT: A K-mer Analysis Toolkit to quality control NGS datasets and genome assemblies

MotivationDe novo assembly of whole genome shotgun (WGS) next-generation sequencing (NGS) data bene[fi]ts from high-quality input with high coverage. However, in practice, determining the quality and quantity of useful reads quickly and in a reference-free manner is not trivial. Gaining a better understanding of the WGS data, and how that data is utilised by assemblers, provides useful insights that can inform the assembly process and result in better assemblies.\n\nResultsWe present the K-mer Analysis Toolkit (KAT): a multi-purpose software toolkit for reference-free quality control (QC) of WGS reads and de novo genome assemblies, primarily via their k-mer frequencies and GC composition. KAT enables users to assess levels of errors, bias and contamination at various stages of the assembly process. In this paper we highlight KATs ability to provide valuable insights into assembly composition and quality of genome assemblies through pairwise comparison of k-mers present in both input reads and the assemblies.\n\nAvailabilityKAT is available under the GPLv3 license at: https://github.com/TGAC/KAT.\n\nContactbernardo.clavijo@earlham.ac.uk\n\nSupplementary InformationSupplementary Information (SI) is available at Bioinformatics online. In addition, the software documentation is available online at: http://kat.readthedocs.io/en/latest/.

Bioinformatics