Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

MetaPalette: A K-mer painting approach for metagenomic taxonomic profiling and quantification of novel strain variation

Metagenomic profiling is challenging in part because of the highly uneven sampling of the tree of life by genome sequencing projects and the limitations imposed by performing phy-logenetic inference at fixed taxonomic ranks. We present the algorithm MetaPalette which uses long k-mer sizes (k = 30, 50) to fit a k-mer \"palette\" of a given sample to the k-mer palette of reference organisms. By modeling the k-mer palettes of unknown organisms, the method also gives an indication of the presence, abundance, and evolutionary relatedness of novel organisms present in the sample. The method returns a traditional, fixed-rank taxonomic profile which is shown on independently simulated data to be one of the most accurate to date. Tree figures are also returned that quantify the relatedness of novel organisms to reference sequences and the accuracy of such figures is demonstrated on simulated spike-ins and a metagenomic soil sample.\n\nThe software implementing MetaPalette is available at: https://github.com/dkoslicki/MetaPalette\n\nPre-trained databases are included for Archaea, Bacteria, Eukaryota, and viruses.

Bioinformatics

Succinct Colored de Bruijn Graphs

Iqbal et al. (Nature Genetics, 2012) introduced the colored de Bruijn graph, a variant of the classic de Bruijn graph, which is aimed at \"detecting and genotyping simple and complex genetic variants in an individual or population\". Because they are intended to be applied to massive population level data, it is essential that the graphs be represented efficiently. Unfortunately, current succinct de Bruijn graph representations are not directly applicable to the colored de Bruijn graph, which require additional information to be succinctly encoded as well as support for non-standard traversal operations. Our data structure dramatically reduces the amount of memory required to store and use the colored de Bruijn graph, with some penalty to runtime, allowing it to be applied in much larger and more ambitious sequence projects than was previously possible.

Bioinformatics

Combating Chagas Disease Through Inhibition of Tiam1, a Rho GTPase Guanine Nucleotide Exchange Factor

Chagas disease is a major cardiovascular affliction primarily endemic to Latin American countries, affecting some ten to twelve million people worldwide. The currently available drugs, Benznidazole and Nifurtimox, are ineffective in the chronic stages and induce severe side effects. In an attempt to improve this situation we use an in silico drug repurposing strategy to correlate drug-protein interactions with positive clinical outcomes. The strategy involves a protein functional site similarity search, along with computational docking studies and, given the findings, a phosphatidylinositol (PIP) strip test to determine the activity of Posaconazole, a recently developed antifungal triazole, in conjunction with Tiam1, a Rho GTPase Guanine Nucleotide Exchange Factor. The results from both computational and in vitro studies indicate possible inhibition of phosphoinositides via Posaconazole, preventing Rho GTPase-induced proliferation of T. cruzi, the etiological agent of Chagas Disease.

Bioinformatics

variancePartition: Interpreting drivers of variation in complex gene expression studies

As genomics studies become more complex and consider multiple sources of biological and technical variation, characterizing these drivers of variation becomes essential to understanding disease biology and regulatory genetics. We describe a statistical and visualization framework, variancePartition, to prioritize drivers of variation with a genome-wide summary, and identify genes that deviate from the genome-wide trend. variancePartition enables rapid interpretation of complex gene expression studies and is applicable to many genomics assays.

Bioinformatics

Reconstruction of ancestral genomes in presence of gene gain and loss.

Since most dramatic genomic changes are caused by genome rearrangements as well as gene duplications and gain/loss events, it becomes crucial to understand their mechanisms and reconstruct ancestral genomes of the given genomes. This problem was shown to be NP-complete even in the \"simplest\" case of three genomes, thus calling for heuristic rather than exact algorithmic solutions. At the same time, a larger number of input genomes may actually simplify the problem in practice as it was earlier illustrated with MGRA, a state-of-the-art software tool for reconstruction of ancestral genomes of multiple genomes.\n\nOne of the key obstacles for MGRA and other similar tools is presence of breakpoint reuses when the same breakpoint region is broken by several different genome rearrangements in the course of evolution. Furthermore, such tools are often limited to genomes composed of the same genes with each gene present in a single copy in every genome. This limitation makes these tools inapplicable for many biological datasets and degrades the resolution of ancestral reconstructions in diverse datasets.\n\nWe address these deficiencies by extending the MGRA algorithm to genomes with unequal gene contents. The developed next-generation tool MGRA2 can handle gene gain/loss events and shares the ability of MGRA to reconstruct ancestral genomes uniquely in the case of limited breakpoint reuse. Furthermore, MGRA2 employs a number of novel heuristics to cope with higher breakpoint reuse and process datasets inaccessible for MGRA. In practical experiments, MGRA2 shows superior performance for simulated and real genomes as compared to other ancestral genomes reconstruction tools. The MGRA2 tool is distributed as an open-source software and can be downloaded from GitHub repository http://github.com/ablab/mgra/. It is also available in the form of a web-server at http://mgra.cblab.org, which makes it readily accessible for inexperienced users.

Bioinformatics

StrainSeeker: fast identification of bacterial strains from unassembled sequencing reads using user-provided guide trees.

BackgroundFast, accurate and high-throughput detection of bacteria is in great demand. The present work was conducted to investigate the possibility of identifying both known and unknown bacterial strains from unassembled next-generation sequencing reads using custom-made guide trees.\n\nResultsA program named StrainSeeker was developed that constructs a list of specific k-mers for each node of any given Newick-format tree and enables rapid identification of bacterial genomes within minutes. StrainSeeker has been tested and shown to successfully identify Escherichia coli strains from mixed samples in less than 5 minutes. StrainSeeker can also identify bacterial strains from highly diverse metagenomics samples. StrainSeeker is available at http://bioinfo.ut.ee/strainseeker.\n\nConclusionsOur novel approach can be useful for both clinical diagnostics and research laboratories because novel bacterial strains are constantly emerging and their fast and accurate detection is very important.

Bioinformatics

AC-PCA: simultaneous dimension reduction and adjustment for confounding variation

Dimension reduction methods are commonly applied to high-throughput biological datasets. However, the results can be hindered by confounding factors, either biologically or technically originated. In this study, we extend Principal Component Analysis to propose AC-PCA for simultaneous dimension reduction and adjustment for confounding variation. We show that AC-PCA can adjust for a) variations across individual donors present in a human brain exon array dataset, and b) variations of different species in a model organism ENCODE RNA-Seq dataset. Our approach is able to recover the anatomical structure of neocortical regions, and to capture the shared variation among species during embryonic development. For gene selection purposes, we extend AC-PCA with sparsity constraints, and propose and implement an efficient algorithm. The methods developed in this paper can also be applied to more general settings.

Bioinformatics

Boiler: Lossy compression of RNA-seq alignments using coverage vectors

We describe Boiler, a new software tool for compressing and querying large collections of RNA-seq alignments. Boiler discards most per-read data, keeping only a genomic coverage vector plus a few empirical distributions summarizing the alignments. Since most per-read data is discarded, storage footprint is often much smaller than that achieved by other compression tools. Despite this, the most relevant per-read data can be recovered; we show that Boiler compression has only a slight negative impact on results given by downstream tools for isoform assembly and quantification. Boiler also allows the user to pose fast and useful queries without decompressing the entire file. Boiler is free open source software available from github.com/jpritt/boiler.

Bioinformatics

CAGEd-oPOSSUM: motif enrichment analysis from CAGE-derived TSSs

SummaryWith the emergence of large-scale Cap Analysis of Gene Expression data sets from individual labs and the FANTOM consortium, one can now analyze the cis-regulatory regions associated with gene transcription at an unprecedented level of refinement. By coupling transcription factor binding site (TFBS) enrichment analysis with CAGE-derived cis-regulatory regions, CAGEd-oPOSSUM can identify TFs that act as key regulators of genes involved in specific mammalian cell and tissue types. The webtool allows for the analysis of CAGE-derived transcription start sites (TSSs) either (i) provided by the user or (ii) selected from ~1,300 mammalian samples from the FANTOM5 project with pre-computed TFBS predicted with JASPAR TF binding profiles. The tool can help power insights into the transcriptional regulation of genes through the study of the specific usage of TSSs within specific cell types and/or under specific conditions.\n\nAvailability and implementationThe CAGEd-oPOSUM web tool is implemented in Perl, MySQL, and Apache and is available at http://cagedop.cmmt.ubc.ca/CAGEd_oPOSSUM. The source code is freely available from GitHub at https://github.com/wassermanlab/CAGEd-oPOSSUM.\n\nContactdave@cmmt.ubc.ca, wyeth@cmmt.ubc.ca, and anthony.mathelier@gmail.com.\n\nSupplementary informationSupplementary Figures are available on the journal website. Supplementary Data is available at http://cagedop.cmmt.ubc.ca/CAGEd_oPOSSUM/archived_results/.

Bioinformatics

BuddySuite: Command-line toolkits for manipulating sequences, alignments, and phylogenetic trees

The ability to manipulate sequence, alignment, and phylogenetic tree files has become an increasingly important skill in the life sciences, whether to generate summary information or to prepare data for further downstream analysis. The command line can be an extremely powerful environment for interacting with these resources, but only if the user has the appropriate general-purpose tools on hand. BuddySuite is a collection of four independent yet interrelated command-line toolkits that facilitate each step in the workflow of sequence discovery, curation, alignment, and phylogenetic reconstruction. Most common sequence, alignment, and tree file formats are automatically detected and parsed, and over 100 tools have been implemented for manipulating these data. The project has been engineered to easily accommodate the addition of new tools, it is written in the popular programming language Python, and is hosted on the Python Package Index and GitHub to maximize accessibility. Documentation for each BuddySuite tool, including usage examples, is available at http://tiny.cc/buddysuite_wiki. All software is open source and freely available through http://research.nhgri.nih.gov/software/BuddySuite.

Bioinformatics

NextflowWorkbench: Reproducible and Reusable Workflows for Beginners and Experts

Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed for beginners, offer a graphical user interface and have been very popular with biologists. In practice, these tools are infrequently used by more experienced bioinformaticians, who may require more flexibility or performance than afforded by the user interfaces, and seem to prefer developing workflows with scripting or command line tools. Here, we present a workflow system, the NextflowWorkbench, which was designed for both beginners and experts, and blends the distinction between user interface and scripting language. This system extends and reuses the popular Nextflow workflow description language and shares its advantages. In contrast to Nextflow, NextflowWorkbench offers an integrated development environment that helps complete beginners get started with workflow development. Auto-completion helps beginners who do not know the syntax of the Nextflow language. Reusable processes provide modular workflows. Programmers will benefit from unique interactive features that help users work more productively with docker containers. We illustrate this tool with a workflow to estimate RNA-Seq counts using Kallisto. The workflow can transparently run either on a laptop computer with docker or on a linux cluster. We found that beginners can be taught how to assemble this workflow in a two hours training session. In conclusion, the NextflowWorkbench simplifies the development of reproducible, implicitly parallel workflows. NextflowWorkbench is distributed under the Apache 2.0 license and available at http://workflow.campagnelab.org.

Bioinformatics

VcfR: an R package to manipulate and visualize VCF format data

Software to call single nucleotide polymorphisms or related genetic variants has converged on the variant call format (VCF) as the output format of choice. This has created a need for tools to work with VCF files. While an increasing number of software exists to read VCF data, many only extract the genotypes without including the data associated with each genotype that describes its quality. We created the R package vcfR to address this issue. We developed a VCF file exploration tool implemented in the R language because R provides an interactive experience and an environment that is commonly used for genetic data analysis. Functions to read and write VCF files into R as well as functions to extract portions of the data and to plot summary statistics of the data are implemented. VcfR further provides the ability to visualize how various parameterizations of the data affect the results. Additional tools are included to integrate sequence (FASTA) and annotation data (GFF) for visualization of genomic regions such as chromosomes. Conversion functions translate data from the vcfR data structure to formats used by other R genetics packages. Computationally intensive functions are implemented in C++ to improve performance. Use of these tools is intended to facilitate VCF data exploration, including intuitive methods for data quality control and easy export to other R packages for further analysis. VcfR thus provides essential, novel tools currently not available in R.

Bioinformatics

Modeling cumulative biological phenomena with Suppes-Bayes causal networks

Several diseases related to cell proliferation are characterized by the accumulation of somatic DNA changes, with respect to wildtype conditions. Cancer and HIV are two common examples of such diseases, where the mutational load in the cancerous/viral population increases over time. In these cases, selective pres sures are often observed along with competition, co-operation and parasitism among distinct cellular clones. Recently, we presented a mathematical framework to model these phenomena, based on a combination of Bayesian inference and Suppes theory of probabilistic causation, depicted in graphical structures dubbed Suppes-Bayes Causal Networks (SBCNs). SBCNs are generative probabilistic graphical models that recapitulate the potential ordering of accumulation of such DNA changes during the progression of the disease. Such models can be inferred from data by exploiting likelihood-based model-selection strategies with regularization. In this paper we discuss the theoretical foun dations of our approach and we investigate in depth the influence on the model-selection task of: (i) the poset based on Suppes theory and (ii) different regulariza tion strategies. Furthermore, we provide an example of application of our framework to HIV genetic data highlighting the valuable insights provided by the inferred SBCN.

Bioinformatics

Diffusion pseudotime robustly reconstructs lineage branching

Single-cell gene expression profiles of differentiating cells encode their intrinsic latent temporal order. We describe an efficient way to robustly estimate this order according to a diffusion pseudotime, which measures transitions on all length scales between cells using diffusion-like random walks. This allows us to identify cells that undergo branching decisions or are in metastable states, and thereby genes differentially regulated at these states.

Bioinformatics

Does the choice of nucleotide substitution models matter topologically?

BackgroundIn the context of a master level programming practical at the computer science department of the Karlsruhe Institute of Technology, we developed and make available an open-source code for testing all 203 possible nucleotide substitution models in the Maximum Likelihood (ML) setting under the common Akaike, corrected Akaike, and Bayesian information criteria. We address the question if model selection matters topologically, that is, if conducting ML inferences under the optimal, instead of a standard General Time Reversible model, yields different tree topologies. We also assess, to which degree models selected and trees inferred under the three standard criteria (AIC, AICc, BIC) differ. Finally, we assess if the definition of the sample size (#sites versus #sites x #taxa) yields different models and, as a consequence, different tree topologies.\n\nResultsWe find that, all three factors (by order of impact: nucleotide model selection, information criterion used, sample size definition) can yield topologically substantially different final tree topologies (topological difference exceeding 10%) for approximately 5% of the tree inferences conducted on the 39 empirical datasets used in our study.\n\nConclusionsWe find that, using the best-fit nucleotide substitution model may change the final ML tree topology compared to an inference under a default GTR model. The effect is less pronounced when comparing distinct information criteria. Nonetheless, in some cases we did obtain substantial topological differences.

Bioinformatics

Genome-Wide Prediction of cis-Regulatory Regions Using Supervised Deep Learning Methods

Identifying active cis-regulatory regions in the human genome is critical for understanding gene regulation and assessing the impact of genetic variation on phenotype. Based on rich data resources such as the Encyclopedia of DNA Elements (ENCODE) and the Functional Annotation of the Mammalian Genome (FANTOM) projects, we introduce DECRES, the first supervised deep learning approach for the identification of enhancer and promoter regions in the human genome. Due to their ability to discover patterns in large and complex data, the introduction of deep learning methods enables a significant advance in our knowledge of the genomic locations of cis-regulatory regions. Using models for well-characterized cell lines, we identify key experimental features that contribute to the predictive performance. Applying DECRES, we delineate locations of 300,000 candidate enhancers genome wide (6.8% of the genome, of which 40,000 are supported by bidirectional transcription data) and 26,000 candidate promoters (0.6% of the genome).

Bioinformatics

Approximate Bayesian bisulphite sequencing analysis

DNA methylation is a key epigenetic modification involved in gene regulation whose contribution to disease susceptibility remains to be fully understood. Here, we present a novel Bayesian smoothing approach (called ABBA) to detect differentially methylated regions (DMRs) from whole-genome bisulphite sequencing (WGBS). We also show how this approach can be leveraged to identify disease-associated changes in DNA methylation, suggesting mechanisms through which these alterations might affect disease. From a data modeling perspective, ABBA has the distinctive feature of automatically adapting to different correlation structures in CpG methylation levels across the genome whilst taking into account the distance between CpG sites as a covariate. Our simulation study shows that ABBA has greater power to detect DMRs than existing methods, providing an accurate identification of DMRs in the large majority of simulated cases. To empirically demonstrate the methods efficacy in generating biological hypotheses, we performed WGBS of primary macrophages derived from an experimental rat system of glomerulonephritis and used ABBA to identify >1,000 disease-associated DMRs. Investigation of these DMRs revealed differential DNA methylation localized to a 600bp region in the promoter of the Ifitm3 gene. This was confirmed by ChIP-seq and RNA-seq analyses, showing differential transcription factor binding at the Ifitm3 promoter by JunD (an established determinant of glomerulonephritis) and a consistent change in Ifitm3 expression. Our ABBA analysis allowed us to propose a new role for Ifitm3 in the pathogenesis of glomerulonephritis via a mechanism involving promoter hypermethylation that is associated with Ifitm3 repression in the rat strain susceptible to glomerulonephritis.

Bioinformatics

SignaFish: a zebrafish-specific signaling pathway resource

Understanding living systems requires an in depth knowledge of the signaling networks that drive cellular homeostasis, regulate intercellular communication and contribute to cell fates during development. Several resources exist to provide high-throughput datasets or manually curated interaction information from human or invertebrate model organisms. We previously developed SignaLink, a uniformly curated, multi-layered signaling resource containing information for human and for the model organisms nematode Caenorhabditis elegans and fruit fly Drosophila melanogaster. Until now, the use of the SignaLink database for zebrafish pathway analysis was limited. To overcome this limitation we created SignaFish (http://signafish.org), a fish-specific signaling resource, built using the concept of SignaLink. SignaFish contains more than 200 curation based signaling interactions, 132 further interactions listed in other resources, and it also lists potential miRNA based regulatory connections for 7 major signaling pathways. From the SignaFish website, users can reach other web resources, such as ZFIN. SignaFish provides signaling or signaling-related interactions that can be examined for each gene, or downloaded for each signaling pathway. We believe that the SignaFish resource will serve as a novel navigating point for experimental design and evaluation for the zebrafish community and for researchers focusing on non-model fish species, such as cyclids.

Bioinformatics