Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Searching and Indexing Genomic Databases via Kernelization

The rapid advance of DNA sequencing technologies has yielded databases of thousands of genomes. To search and index these databases effectively, it is important that we take advantage of the similarity between those genomes. Several authors have recently suggested searching or indexing only one reference genome and the parts of the other genomes where they differ. In this paper we survey the twenty-year history of this idea and discuss its relation to kernelization in parameterized complexity.

Bioinformatics

DensiTree 2: Seeing Trees Through the Forest

MotivationPhylogenetic analysis like Bayesian MCMC or bootstrapping result in a collection of trees. Trees are discrete objects and it is generally difficult to get a mental grip on a distributions over trees. Visualisation tools like DensiTree can give good intuition on tree distributions. It works by drawing all trees in the set transparently thus highlighting areas where the tree in the set agrees. In this way, both uncertainty in clade heights and uncertainty in topology can be visualised. In our experience, a vanilla DensiTree can turn out to be misleading in that it shows too much uncertainty due to wrongly ordering taxa or due to unlucky placement of internal nodes.\n\nResultsDensiTree is extended to allow visualisation of meta-data associated with branches such as population size and evolutionary rates. Furthermore, geographic locations of taxa can be shown on a map, making it easy to visually check there is some geographic pattern in a phylogeny. Taxa orderings have a large impact on the layout of the tree set, and advances have been made in finding better orderings resulting in significantly more informative visualisations. We also explored various methods for positioning internal nodes, which can improve the quality of the image. Together, these advances make it easier to comprehend distributions over trees.\n\nAvailabilityDensiTree is freely available from http://compevol.auckland.ac.nz/software/.

Bioinformatics

CIDER: a pipeline for detecting waves of coordinated transcriptional regulation in gene expression time-course data

Cell adaptability to environmental changes is conferred by complex transcriptional regulatory networks, which respond to external stimuli by modulating the expression dynamics of each gene. Hence, deciphering the network of transcriptional regulation is remarkably important, but proves to be extremely challenging, mainly due to the unfavorable ratio between the number of available observations and the number of parameters to estimate. Most of the existing computational methods for the inference of transcriptional networks consider steady-state gene expression datasets, and produce models of transcriptional regulation best explaining the observed static gene expression.\n\nGene expression time-courses are an emergent typology of gene expression data, paving the way to the characterization of the time-dependent dynamics of transcriptional regulation.\n\nIn this work we introduce the Complexity Invariant Dynamic Time Warping motif EnRichment (CIDER) analysis, a novel computational pipeline to identify the prominent waves of coordinated gene transcription induced in cells by external stimuli, and determine which TFs are involved in the coordination of gene transcription. The CIDER pipeline combines unsupervised time series clustering and motif enrichment analysis to first detect transcriptional expression patterns, and then identify the TFs over-represented in the promoter regions of gene sets with similar expression dynamics.\n\nThe ability of CIDER to correctly identify regulatory interactions is assessed on a realistic synthetic dataset of gene expression time-courses, generated by simulating the effects of knock-out perturbations on the E. coli regulatory network.\n\nThe CIDER source code and the validation datasets are available on request from the corresponding author.

Bioinformatics

HISAT: Hierarchical Indexing for Spliced Alignment of Transcripts

HISAT is a new, highly efficient system for alignment of sequences from RNA sequencing experiments that achieves dramatically faster performance than previous methods. HISAT uses a new indexing scheme, hierarchical indexing, which is based on the Burrows-Wheeler transform and the Ferragina-Manzini (FM) index. Hierarchical indexing employs two types of indexes for alignment: (1) a whole-genome FM index to anchor each alignment, and (2) numerous local FM indexes for very rapid extensions of these alignments. HISATs hierarchical index for the human genome contains 48,000 local FM indexes, each representing a genomic region of ~64,000 bp. The algorithm includes several customized alignment strategies specifically designed for mapping RNA-seq reads across multiple exons. In tests on a variety of real and simulated data sets, we show that HISAT is the fastest system currently available, approximately 50 times faster than TopHat2 and 12 times faster than GSNAP, with equal or better accuracy than any other method. Despite its very large number of indexes, HISAT requires only 4.3 Gigabytes of memory to align reads to the human genome. HISAT supports genomes of any size, including those larger than 4 billion bases. HISAT is available as free, open-source software from http://www.ccb.jhu.edu/software/hisat.

Bioinformatics

BioWardrobe: an integrated platform for analysis of epigenomics and transcriptomics data

BioWardrobe is an integrated platform that empowers users to store, visualize and analyze epigenomics and transcriptomics data, without the need for programming expertise, by using a biologist-friendly webbased user interface. Predefined pipelines automate download of data from core facilities or public databases, calculate RPKMs and identify peaks and display results via the interactive web interface. Additional capabilities include analyzing differential gene expression and DNA-protein binding and creating average tag density profiles and heatmaps.

Bioinformatics

Shrinkage of dispersion parameters in the binomial family, with application to differential exon skipping

The prevalence of sequencing experiments in genomics has led to an increased use of methods for count data in analyzing high-throughput genomic data to perform analyses. The importance of shrinkage methods in improving the performance of statistical methods remains. A common example is that of gene expression data, where the counts per gene are often modeled as some form of an over-dispersed Poisson. In this case, shrinkage estimates of the per-gene dispersion parameter have led to improved estimation of dispersion in the case of a small number of samples.\n\nWe address a different count setting introduced by the use of sequencing data: comparing differential proportional usage via an over-dispersed binomial model. This is motivated by our interest in testing for differential exon skipping in mRNA-Seq experiments. We introduce a novel method that is developed by modeling the dispersion based on the double binomial distribution proposed by Efron (1986). Our method (WEB-Seq) is an empirical bayes strategy for producing a shrunken estimate of dispersion and effectively detects differential proportional usage, and has close ties to the weighted-likelihood strategy of edgeR developed for gene expression data (Robinson and Smyth, 2007; Robinson et al., 2010). We analyze its behavior on simulated data sets as well as real data and show that our method is fast, powerful and gives accurate control of the FDR compared to alternative approaches. We provide implementation of our methods in the R package DoubleExpSeq available on CRAN.

Bioinformatics

ViennaNGS: A toolbox for building efficient next-generation sequencing analysis pipelines

Recent achievements in next-generation sequencing (NGS) technologies lead to a high demand for reuseable software components to easily compile customized analysis workflows for big genomics data. We present ViennaNGS, an integrated collection of Perl modules focused on building efficient pipelines for NGS data processing. It comes with functionality for extracting and converting features from common NGS file formats, computation and evaluation of read mapping statistics, as well as normalization of RNA abundance. Moreover, ViennaNGS provides software components for identification and characterization of splice junctions from RNA-seq data, parsing and condensing sequence motif data, automated construction of Assembly and Track Hubs for the UCSC genome browser, as well as wrapper routines for a set of commonly used NGS command line tools.

Bioinformatics

FORGE : A tool to discover cell specific enrichments of GWAS associated SNPs in regulatory regions.

Genome wide association studies provide an unbiased discovery mechanism for numerous human diseases. However, a frustration in the analysis of GWAS is that the majority of variants discovered do not directly alter protein-coding genes. We have developed a simple analysis approach that detects the tissue-specific regulatory component of a set of GWAS SNPs by identifying enrichment of overlap with DNase I hotspots from diverse tissue samples. Functional element Overlap analysis of the Results of GWAS Experiments (FORGE) is available as a web tool and as standalone software and provides tabular and graphical summaries of the enrichments. Conducting FORGE analysis on SNP sets for 260 phenotypes available from the GWAS catalogue reveals numerous overlap enrichments with tissue-specific components reflecting the known aetiology of the phenotypes as well as revealing other unforeseen tissue involvements that may lead to mechanistic insights for disease.

Bioinformatics

RAX2: genome-wide detection of condition-associated transcription variation

Almost all of mammalian genes have mRNA variants due to alternative promoters, alternative splice sites, and alternative cleavage and polyadenylation sites. In most cases, change in transcript due to choosing alternative cleavage and polyadenylation sites does not lead to change in protein sequence, while selection of alternative promoters and alternative splice sites would alter protein sequence. Nevertheless, all these alternations would give rise to different RNA isoforms. Selection of alternative RNA isoforms has been found to be associated with change in condition. For example, many studies have revealed that alternative cleavage and polyadenylation (poly(A)) are correlated to proliferation, differentiation, and cellular transformation. Thus, unlike gene expression in microarray, change in usage of splice sites or poly(A) sites associated with conditions does not involve differential expression and hence cannot be detected by differential analysis methods but can be done by association methods. Traditional association methods such as Pearson chi-square test and Fisher Exact test are single test methods and do not work on the RNA count data derived from replicate libraries. For this reason, we here developed a large-scale association method, called ranking analysis of chi-squares (RAX2). Simulations demonstrated that RAX2 worked well for finding association of changes in usage of poly(A) sites with condition change. We applied our RAX2 to our primary T-cell transcriptomic data of over 9899 tags scattered in 3812 genes and found that 1610 (16.3%) tags were associated in transcription with immune stimulation at FDR < 0.05 and most of these tags associated with stimulation also had differential expression. Analysis of two and three tags within genes revealed that under immune stimulation, short RNA isoforms were significantly preferably used. Like cell proliferation and division, short RNA isoforms are highly prioritized to be used for cell growth.

Bioinformatics

When Less is More: "Slicing" Sequencing Data Improves Read Decoding Accuracy and De Novo Assembly Quality

Since the invention of DNA sequencing in the seventies, computational biologists have had to deal with the problem de novo genome assembly with limited (or insufficient) depth of sequencing. In this work, for the first time we investigate the opposite problem, that is, the challenge of dealing with excessive depth of sequencing. Specifically, we explore the effect of ultra-deep sequencing data in two domains: (i) the problem of decoding reads to BAC clones (in the context of the combinatorial pooling design proposed in [1]), and (ii) the problem of de novo assembly of BAC clones. Using real ultra-deep sequencing data, we show that when the depth of sequencing increases over a certain threshold, sequencing errors make these two problems harder and harder (instead of easier, as one would expect with error-free data), and as a consequence the quality of the solution degrades with more and more data. For the first problem, we propose an effective solution based on \"divide and conquer\": we \"slice\" a large dataset into smaller samples of optimal size, decode each slice independently, then merge the results. Experimental results on over 15,000 barley BACs and over 4,000 cowpea BACs demonstrate a significant improvement in the quality of the decoding and the final assembly. For the second problem, we show for the first time that modern de novo assemblers cannot take advantage of ultra-deep sequencing data.

Bioinformatics

SCIE: Information Extraction for Spinal Cord Injury Preclinical Experiments - A Webservice and Open Source Toolkit

Translational neuroscience in the field of spinal cord injuries (SCI) faces a strong disproportion between immense preclinical research efforts and a lack of therapeutic approaches successful in human patients: Currently, preclinical research on SCI yields more than 3,000 new publications per year (8,000 when including the whole central nervous system, growing at an exponential rate), whereas none of the resulting therapeutic concepts has led to functional recovery of neural tissue in humans. Improving clinical researchers information access therefore carries the potential to support more effective selection of promising therapy candidates from preclinical studies. Thus, automated information extraction from scientific publications contributes to enabling meta studies and therapy grading by aggregating relevant information from the entire body of previous work on SCI.\n\nWe present SCIE, an automated information extraction pipeline capable of detecting relevant information in SCI publications based on ontological entity and probabilistic relation detection. The input are plain text or PDF documents. As output, the user choses between an online visualization or a machine-readable format. Compared to human gold standard annotations, our system achieves an average extraction performance of 76 % precision and 52 % recall (F1-measure 0.59).\n\nAn instance of the webservice is available at http://scie.sc.cit-ec.uni-bielefeld.de/. SCIE is free software licensed under the AGPL and can be downloaded for local installation at http://opensource.cit-ec.de/projects/scie/.

Bioinformatics

Oxford Nanopore Sequencing, Hybrid Error Correction, and de novo Assembly of a Eukaryotic Genome

Monitoring the progress of DNA molecules through a membrane pore has been postulated as a method for sequencing DNA for several decades. Recently, a nanopore-based sequencing instrument, the Oxford Nanopore MinION, has become available that we used for sequencing the S. cerevisiae genome. To make use of these data, we developed a novel open-source hybrid error correction algorithm Nanocorr (https://github.com/jgurtowski/nanocorr) specifically for Oxford Nanopore reads, as existing packages were incapable of assembling the long read lengths (5-50kbp) at such high error rate (between [~]5 and 40% error). With this new method we were able to perform a hybrid error correction of the nanopore reads using complementary MiSeq data and produce a de novo assembly that is highly contiguous and accurate: the contig N50 length is more than ten-times greater than an Illumina-only assembly (678kb versus 59.9kbp), and has greater than 99.88% consensus identity when compared to the reference. Furthermore, the assembly with the long nanopore reads presents a much more complete representation of the features of the genome and correctly assembles gene cassettes, rRNAs, transposable elements, and other genomic features that were almost entirely absent in the Illumina-only assembly.\n\nReviewer link to datahttp://schatzlab.cshl.edu/data/nanocorr/

Bioinformatics

Software for the analysis and visualization of deep mutational scanning data

BackgroundDeep mutational scanning is a technique to estimate the impacts of mutations on a gene by using deep sequencing to count mutations in a library of variants before and after imposing a functional selection. The impacts of mutations must be inferred from changes in their counts after selection.\n\nResultsI describe a software package, dms_tools, to infer the impacts of mutations from deep mutational scanning data using a likelihood-based treatment of the mutation counts. I show that dms_tools yields more accurate inferences on simulated data than simply calculating ratios of counts pre-and post-selection. Using dms_tools, one can infer the preference of each site for each amino acid given a single selection pressure, or assess the extent to which these preferences change under different selection pressures. The preferences and their changes can be intuitively visualized with sequence-logo-style plots created using an extension to weblogo.\n\nConclusionsdms_tools implements a statistically principled approach for the analysis and subsequent visualization of deep mutational scanning data.

Bioinformatics

JAFFA: High sensitivity transcriptome-focused fusion gene detection.

Genomic instability is a hallmark of cancer and, as such, structural alterations and fusion genes are common events in the cancer landscape. RNA sequencing (RNA-Seq) is a powerful method for profiling cancers, but current methods for identifying fusion genes are optimized for short reads. JAFFA (https://code.google.com/p/jaffa-project/) is a sensitive fusion detection method that clearly out-performs other methods with reads of 100bp or greater. JAFFA compares a cancer transcriptome to the reference transcriptome, rather than the genome, where the cancer transcriptome is inferred using long reads directly or by de novo assembling short reads.

Bioinformatics

SW#db: GPU-accelerated exact sequence similarity database search

The deluge of next-generation sequencing (NGS) data and expanding database poses higher requirements for protein similarity search. State-of-the-art tools such as BLAST are not fast enough to cope with these requirements. Because of that it is necessary to create new algorithms that will be faster while keeping similar sensitivity levels. The majority of protein similarity search methods are based on a seed-and-extend approach which uses standard dynamic programming algorithms in the extend phase. In this paper we present a SW#db tool and library for exact similarity search. Although its running times, as standalone tool, are comparable to running times of BLAST it is primarily designed for the extend phase where there are reduced number of candidates in the database. It uses both GPU and CPU parallelization and when we measured multiple queries on Swiss-prot and Uniref90 databases SW#db was 4 time faster than SSEARCH, 6-10 times faster than CUDASW++ and more than 20 times faster than SSW.

Bioinformatics

The LUX Score: A Metric for Lipidome Homology

MotivationWe propose a method for estimating lipidome homologies analogous to the ones used in sequence analysis and phylo-genetics.\n\nResultsAlgorithms were developed to quantify the structural similarity between lipids and to compute chemical space models of sets of lipid structures and lipidomes. When all lipid molecules of the LIPIDMAPS structure database were mapped in such a chemical space, they automatically formed clusters corresponding to conventional chemical families. Homologies between the lipidomes of four yeast strains based on our LUX score reflected the genetic relationship, although the score is based solely on lipid structures.\n\nAvailabilitywww.lux.fz-borstel.de

Bioinformatics

Phylesystem: a git-based data store for community curated phylogenetic estimates

1 MotivationPhylogenetic estimates from published studies can be archived using general platforms like Dryad (Vision, 2010) or TreeBASE (Sanderson et al., 1994). Such services fulfill a crucial role in ensuring transparency and reproducibility in phylogenetic research. However, digital tree data files often require some editing (e.g. rerooting) to improve the accuracy and reusability of the phylogenetic statements. Furthermore, establishing the mapping between tip labels used in a tree and taxa in a single common taxonomy dramatically improves the ability of other researchers to reuse phylogenetic estimates. Because the process of curating a published phylogenetic estimate is not error-free, retaining a full record of the provenance of edits to a tree is crucial for openness, allowing editors to receive credit for their work, and making errors introduced during curation easier to correct.\n\n2 ResultsHere we report the development of software infrastructure to support the open curation of phylogenetic data by the community of biologists. The backend of the system provides an interface for the standard database operations of creating, reading, updating, and deleting records by making commits to a git repository. The record of the history of edits to a tree is preserved by gits version control features. Hosting this data store on GitHub (2014) provides open access to the data store using tools familiar to many developers. We have deployed a server running the \"phylesystem-api\", which wraps the interactions with git and GitHub. The Open Tree of Life project has also developed and deployed a JavaScript application that uses the phylesystem-api and other web services to enable input and curation of published phylogenetic statements.\n\n3 AvailabilitySource code for the web service layer is available at https://github.com/OpenTreeOfLife/phylesystem-api. The data store can be cloned from: https://github.com/OpenTreeOfLife/phylesystem. A web application that uses the phylesystem web services is deployed at https://tree.opentreeoflife.org/curator. Code for that tool is available from https://github.com/OpenTreeOfLife/opentree\n\n4 Contactmtholder@gmail.com

Bioinformatics

MultiMeta: an R package for meta-analysing multi-phenotype genome-wide association studies

SummaryAs new methods for multivariate analysis of Genome Wide Association Studies (GWAS) become available, it is important to be able to combine results from different cohorts in a meta-analysis. The R package MultiMeta provides an implementation of the inverse-variance based method for meta-analysis, generalized to an n-dimensional setting.\n\nAvailabilityThe R package MultiMeta can be downloaded from CRAN Contact: dragana.vuckovic@burlo.trieste.it

Bioinformatics