Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Discrete Distributional Differential Expression (D3E) - A Tool for Gene Expression Analysis of Single-cell RNA-seq Data

The advent of high throughput RNA-seq at the single-cell level has opened up new opportunities to elucidate the heterogeneity of gene expression. One of the most widespread applications of RNA-seq is to identify genes which are differentially expressed between two experimental conditions. Here, we present a discrete, distributional method for differential gene expression (D3E), a novel algorithm specifically designed for single-cell RNA-seq data. We use synthetic data to evaluate D3E, demonstrating that it can detect changes in expression, even when the mean level remains unchanged. Since D3E is based on an analytically tractable stochastic model, it provides additional biological insights by quantifying biologically meaningful properties, such as the average burst size and frequency. We use D3E to investigate experimental data, and with the help of the underlying model, we directly test hypotheses about the driving mechanism behind changes in gene expression.

Bioinformatics

Excess False Positive Rates in Methods for Differential Gene Expression Analysis using RNA-Seq Data

MotivationAn important property of a valid method for testing for differential expression is that the false positive rate should at least roughly correspond to the p-value cutoff, so that if 10,000 genes are tested at a p-value cutoff of 10-4, and if all the null hypotheses are true, then there should be only about 1 gene declared to be significantly differentially expressed. We tested this by resampling from existing RNA-Seq data sets and also by matched negative binomial simulations.\n\nResultsMethods we examined, which rely strongly on a negative binomial model, such as edgeR, DESeq, and DESeq2, show large numbers of false positives in both the resampled real-data case and in the simulated negative binomial case. This also occurs with a negative binomial generalized linear model function in R. Methods that use only the variance function, such as limma-voom, do not show excessive false positives, as is also the case with a variance stabilizing transformation followed by linear model analysis with limma. The excess false positives are likely caused by apparently small biases in estimation of negative binomial dispersion and, perhaps surprisingly, occur mostly when the mean and/or the dispersion is high, rather than for low-count genes.\n\nContactdmrocke@ucdavis.edu, lruan@ucdavis.edu, yilzhang@ucdavis.edu, gt4636b@gatech.edu, bpdur-bin@ucdavis.edu, saviran@ucdavis.edu.\n\nSupplementary InformationThe computational tools developed for this study are freely available via our website http://dmrocke.ucdavis.edu/software.html. They can be downloaded as R code or run directly through an interactive web-based shiny application to reproduce the analysis presented here per a users choice of dataset and the methods to be evaluated.

Bioinformatics

bModelTest: Bayesian phylogenetic site model averaging and model comparison

BackgroundReconstructing phylogenies through Bayesian methods has many benefits, which include providing a mathematically sound framework, providing realistic estimates of uncertainty and being able to incorporate different sources of information based on formal principles. Bayesian phylogenetic analyses are popular for interpreting nucleotide sequence data, however for such studies one needs to specify a site model and associated substitution model. Often, the parameters of the site model is of no interest and an ad-hoc or additional likelihood based analysis is used to select a single site model.\n\nResultsbModelTest allows for a Bayesian approach to inferring and marginalizing site models in a phylogenetic analysis. It is based on trans-dimensional Markov chain Monte Carlo (MCMC) proposals that allow switching between substitution models as well as estimating the posterior probability for gamma-distributed rate heterogeneity a proportion of invariable sites and unequal base frequencies. The model can be used with the full set of time-reversible models on nucleotides, but we also introduce and demonstrate the use of two subsets of time-reversible substitution models.\n\nConclusionWith the new method the site model can be inferred (and marginalized) during the MCMC analysis and does not need to be pre-determined, as is now often the case in practice, by likelihood-based methods. The method is implemented in the bModelTest package of the popular BEAST 2 software, which is open source, licensed under the GNU Lesser General Public License and allows joint site model and tree inference under a wide range of models.

Bioinformatics

MAST: A flexible statistical framework for assessing transcriptional changes and characterizing heterogeneity in single-cell RNA-seq data.

Single-cell transcriptomic profiling enables the unprecedented interrogation of gene expression heterogeneity in rare cell populations that would otherwise be obscured in bulk RNA sequencing experiments. The stochastic nature of transcription is revealed in the bimodality of single-cell transcriptomic data, a feature shared across single-cell expression platforms. There is, however, a paucity of computational tools that take advantage of this unique characteristic. We present a new methodology to analyze single-cell transcriptomic data that models this bimodality within a coherent generalized linear modeling framework. We propose a two-part, generalized linear model that allows one to characterize biological changes in the proportions of cells that are expressing each gene, and in the positive mean expression level of that gene. We introduce the cellular detection rate, the fraction of genes turned on in a cell, and show how it can be used to simultaneously adjust for technical variation and so-called \"extrinsic noise\" at the single-cell level without the use of control genes. Our model permits direct inference on statistics formed by collections of genes, facilitating gene set enrichment analysis. The residuals defined by such models can be manipulated to interrogate cellular heterogeneity and gene-gene correlation across cells and conditions, providing insights into the temporal evolution of networks of co-expressed genes at the single-cell level. Using two single-cell RNA-seq datasets, including newly generated data from Mucosal Associated Invariant T (MAIT) cells, we show how model residuals can be used to identify significant changes across biologically relevant gene sets that are missed by other methods and characterize cellular heterogeneity in response to stimulation.

Bioinformatics

Automated Contamination Detection in Single-Cell Sequencing

Novel methods for the sequencing of single-cell DNA offer tremendous opportunities. However, many techniques are still in their infancy and a major obstacle is given by sample contamination with foreign DNA. In this contribution, we present a pipeline that allows for fast, automated detection of contaminated samples by the use of modern machine learning methods. First, a vectorial representation of the genomic data is obtained using oligonucleotide signatures. Using non-linear subspace projections, data is transformed to be suitable for automatic clustering. This allows for the detection of one vs. more genomes (clusters) in a sample. As clustering is an ill-posed problem, the pipeline relies on a thorough choice of all involved methods and parameters. We give an overview of the problem and evaluate techniques suitable for this task.

Bioinformatics

Overcoming analytical reliability issues in clinical proteomics using rank-based network approaches

Proteomics is poised to play critical roles in clinical research. However, due to limited coverage and high noise, integration with powerful analysis algorithms is necessary. In particular, network-based algorithms can improve selection of reproducible features in spite of incomplete proteome coverage, technical inconsistency or high inter-sample variability. We define analytical reliability on three benchmarks --- precision/recall rates, feature-selection stability and cross-validation accuracy. Using these, we demonstrate the insufficiencies of commonly used Students t-test and Hypergeometric enrichment. Given advances in sample sizes, quantitation accuracy and coverage, we are now able to introduce and evaluate Ranked-Based Network Approaches (RBNAs) for the first time in proteomics. These include SNET (SubNETwork), FSNET (FuzzySNET), PFSNET (PairedFSNET). We also introduce for the first time, PPFSNET(samplePairedPFSNET), which is a paired-sample variant of PFSNET. RBNAs (particularly PFSNET and PPFSNET) excelled on all three benchmarks and can make consistent and reproducible predictions even in the small-sample size scenario (n=4). Given these qualities, RBNAs represent an important advancement in network biology, and is expected to see practical usage, particularly in clinical biomarker and drug target prediction.

Bioinformatics

Tools and pipelines for BioNano data: molecule assembly pipeline and FASTA super scaffolding tool

BackgroundGenome assembly remains an unsolved problem. Assembly projects face a range of hurdles that confound assembly. Thus a variety of tools and approaches are needed to improve draft genomes.\n\nResultsWe used a custom assembly workflow to optimize consensus genome map assembly, resulting in an assembly equal to the estimated length of the Tribolium castaneum genome and with an N50 of more than 1 Mb. We used this map for super scaffolding the T. castaneum sequence assembly, more than tripling its N50 with the program Stitch.\n\nConclusionsIn this article we present software that leverages consensus genome maps assembled from extremely long single molecule maps to increase the contiguity of sequence assemblies. We report the results of applying these tools to validate and improve a 7x Sanger draft of the T. castaneum genome.

Bioinformatics

Sequence evidence for common ancestry of eukaryotic endomembrane coatomers

Eukaryotic cells are defined by compartments through which the trafficking of macromolecules is mediated by large complexes, such as the nuclear pore, transport vesicles and intraflagellar transport. The assembly and maintenance of these complexes is facilitated by endomembrane coatomers, long suspected to be divergently related on the basis of structural and more recently phylogenomic analysis. By performing supervised walks in sequence space across coatomer superfamilies, we uncover subtle sequence patterns that have remained elusive to date, ultimately unifying eukaryotic coatomers by divergent evolution. The conserved residues shared by 3,502 endomembrane coatomer components are mapped onto the solenoid superhelix of nucleoporin and COPII protein structures, thus determining the invariant elements of coatomer architecture. This ancient structural motif can be considered as a universal signature connecting eukaryotic coatomers involved in multiple cellular processes across cell physiology and human disease.

Bioinformatics

EVfold.org: Evolutionary Couplings and Protein 3D Structure Prediction

Recently developed maximum entropy methods infer evolutionary constraints on protein function and structure from the millions of protein sequences available in genomic databases. The EVfold web server (at EVfold.org) makes these methods available to predict functional and structural interactions in proteins. The key algorithmic development has been to disentangle direct and indirect residue-residue correlations in large multiple sequence alignments and derive direct residue-residue evolutionary couplings (EVcouplings or ECs). For proteins of unknown structure, distance constraints obtained from evolutionarily couplings between residue pairs are used to de novo predict all-atom 3D structures, often to good accuracy. Given sufficient sequence information in a protein family, this is a major advance toward solving the problem of computing the native 3D fold of proteins from sequence information alone.\n\nAvailabilityEVfold server at http://evfold.org/\n\nContactevfoldtest@gmail.com\n\nAbbreviations

Bioinformatics

Biographika: rich interactive data visualizations on the web for the research community

When visualizing scientific data one of the current bottlenecks is the lack of interactivity. There already exist many options to build static data visualizations such as R, Matlab or Microsoft Excel among others. On the other hand, we can also find many different pieces of software with a broader or more specific aim that however must be installed locally. There is therefore a gap that is not covered by any of these latter two worlds. Here is where Biographika comes in handy since it provides scientists in general, and more specifically bioinformaticians, with a way of being able to use interactive rich visualizations on the web as part of their daily research. This first version of Biographika includes a set of charts combining diverse approaches that are thought to give users different perspectives of their research data. For the sake of interoperability and expressivity among other reasons we are using D3.js [1], the de facto standard visualization JavaScript library for manipulating documents based on data. But not only that, we incorporate new approaches as the fact of having fully interactive 3D charts that can be easily integrated with the rest of visualizations; providing the possibility of analyzing multidimensional data in a way that could otherwise be difficult to tackle. For that we use X3DOM[2], an open-source framework and runtime for 3D graphics on the Web that eases the integration of HTML5 and declarative 3D content. And last but not least, Biographika is also conceived as an effort to provide the data visualization layer for Bio4j [3] that researchers have been lacking in the past few years.

Bioinformatics

PaxtoolsR: Pathway Analysis in R Using Pathway Commons

PurposePaxtoolsR package enables access to pathway data represented in the BioPAX format and made available through the Pathway Commons webservice for users of the R language. Features include the extraction, merging, and validation of pathway data represented in the BioPAX format. This package also provides novel pathway datasets and advanced querying features for R users through the Pathway Commons webservice allowing users to query, extract, and retrieve data and integrate this data with local BioPAX datasets.\n\nAvailabilityThe PaxtoolsR package is compatible with R 3.1.1 on Windows, Mac OS X, and Linux using Bioconductor 3.0 and is available through the Bioconductor R package repository along with source code and a tutorial vignette describing common tasks, such as data visualization and gene set enrichment analysis. Source code and documentation are at http://bioconductor.org/packages/release/bioc/html/paxtoolsr.html. This plugin is free, open-source and licensed under the GNU Lesser General Public License (LGPL) v3.0.\n\nContactpaxtools@cbio.mskcc.org

Bioinformatics

TESS: Bayesian inference of lineage diversification rates from (incompletely sampled) molecular phylogenies in R

SummaryMany fundamental questions in evolutionary biology entail estimating rates of lineage diversification (speciation - extinction). We develop a flexible Bayesian framework for specifying an effectively infinite array of diversification models--where rates are constant, vary continuously, or change episodically through time--and implement numerical methods to estimate parameters of these models from molecular phylogenies, even when species sampling is incomplete. Additionally we provide robust methods for comparing the relative and absolute fit of competing branching-process models to a given tree, thereby providing rigorous tests of biological hypotheses regarding patterns and processes of lineage diversification.\n\nAvailability and implementationthe source code for TESS is freely available at http://cran.r-project.org/web/packages/TESS/.\n\nContactSebastian.Hoehna@gmail.com

Bioinformatics

Inference under a Wright-Fisher model using an accurate beta approximation

The large amount and high quality of genomic data available today enables, in principle, accurate inference of evolutionary history of observed populations. The Wright-Fisher model is one of the most widely used models for this purpose. It describes the stochastic behavior in time of allele frequencies and the influence of evolutionary pressures, such as mutation and selection. Despite its simple mathematical formulation, exact results for the distribution of allele frequency (DAF) as a function of time are not available in closed analytic form. Existing approximations build on the computationally intensive diffusion limit, or rely on matching moments of the DAF. One of the moment-based approximations relies on the beta distribution, which can accurately describe the DAF when the allele frequency is not close to the boundaries (zero and one). Nonetheless, under a Wright-Fisher model, the probability of being on the boundary can be positive, corresponding to the allele being either lost or fixed. Here, we introduce the beta with spikes, an extension of the beta approximation, which explicitly models the loss and fixation probabilities as two spikes at the boundaries. We show that the addition of spikes greatly improves the quality of the approximation. We additionally illustrate, using both simulated and real data, how the beta with spikes can be used for inference of divergence times between populations, with comparable performance to existing state-of-the-art method.

Bioinformatics

Discovery of methylation loci and analyses of differential methylation from replicated high-throughput sequencing data

Cytosine methylation is widespread in most eukaryotic genomes and is known to play a substantial role in various regulatory pathways. Unmethylated cytosines may be converted to uracil through the addition of sodium bisulphite, allowing genome-wide quantification of cytosine methylation via high-throughput sequencing. The data thus acquired allows the discovery of methylation loci; contiguous regions of methylation consistently methylated across biological replicates. The mapping of these loci allows for associations with other genomic factors to be identified, and for analyses of differential methylation to take place.\n\nThe segmentSeq R package is extended to identify methylation loci from high-throughput sequencing data from multiple experimental conditions. A statistical model is then developed that accounts for biological replication and variable rates of non-conversion of cytosines in each sample to compute posterior likelihoods of methylation at each locus within an empirical Bayesian framework. The same model is used as a basis for analysis of differential methylation between multiple experimental conditions with the baySeq R package. We demonstrate this method through an analysis of data derived from Dicer-like mutants in Arabidopsis that reveals complex interactions between the different Dicer-like mutants and their methylation pathways. We also show in simulation studies that this approach can be significantly more powerful in the detection of differential methylation than existing methods.

Bioinformatics

On enhancing variation detection through pan-genome indexing

Detection of genomic variants is commonly conducted by aligning a set of reads sequenced from an individual to the reference genome of the species and analyzing the resulting read pileup. Typically, this process finds a subset of variants already reported in databases and additional novel variants characteristic to the sequenced individual. Most of the effort in the literature has been put to the alignment problem on a single reference sequence, although our gathered knowledge on species such as human is pan-genomic: We know most of the common variation in addition to the reference sequence. There have been some efforts to exploit pan-genome indexing, where the most widely adopted approach is to build an index structure on a set of reference sequences containing observed variation combinations.\n\nThe enhancement in alignment accuracy when using pan-genome indexing has been demonstrated in experiments, but so far the above multiple references pan-genome indexing approach has not been tested on its final goal, that is, in enhancing variation detection. This is the focus of this article: We study a generic approach to add variation detection support on top of the multiple references pan-genomic indexing approach. Namely, we study the read pileup on a multiple alignment of reference genomes, and propose a heaviest path algorithm to extract a new recombined reference sequence. This recombined reference sequence can then be utilized in any standard read alignment and variation detection workflow. We demonstrate that the approach enhances variation detection on realistic data sets.

Bioinformatics

SSCM: A method to analyze and predict the pathogenicity of sequence variants

AbstractThe advent of cost-effective DNA sequencing has provided clinics with high-resolution information about patients genetic variants, which has resulted in the need for efficient interpretation of this genomic data. Traditionally, variant interpretation has been dominated by many manual, time-consuming processes due to the disparate forms of relevant information in clinical databases and literature. Computational techniques promise to automate much of this, and while they currently play only a supporting role, their continued improvement for variant interpretation is necessary to tackle the problem of scaling genetic sequencing to ever larger populations. Here, we present SSCM-Pathogenic, a genome-wide, allele-specific score for predicting variant pathogenicity. The score, generated by a semi-supervised clustering algorithm, shows predictive power on clinically relevant mutations, while also displaying predictive ability in noncoding regions of the genome.

Bioinformatics

In Silico Predictive Modeling of CRISPR/Cas9 guide efficiency

The CRISPR/Cas9 system provides unprecedented genome editing capabilities; however, several facets of this system are under investigation for further characterization and optimization, including the choice of guide RNA that directs Cas9 to target DNA. In particular, given that one would like to target the protein-coding region of a gene, hundreds of guides satisfy the basic constraints of the CRISPR/Cas9 Protospacer Adjacent Motif sequence (PAM); however, not all of these guides actually generate gene knockouts with equal efficiency. Leveraging a broad set of experimental measurements of guide knockout efficiency, we introduce a state-of-the art in silico modeling approach to identify guides that will lead to more effective gene knockout. We first investigated which guide and gene features are critical for prediction (e.g., single- and di-nucleotide identity of the gene target), which are helpful (e.g., thermodynamics), and which are predictive but redundant (e.g., microhomology). We also investigated evaluation measures for comparing predictive models in the present context, suggesting that Area Under the Receiver Operating Curve is not ideal. Finally, we explored a variety of different model classes and found that use of gradient-boosted regression trees produced the best predictive performance. Pointers to our open-source software, code, and prediction server will be available at http://research.microsoft.com/en-us/projects/azimuth.

Bioinformatics

Modeling small RNA competition in C. elegans

Website SummarySmall RNAs are important regulators of gene expression; however, the relationship between small RNAs is poorly understood. Studying the crosstalk between small RNA pathways can help understand how gene expression is regulated. One hypothesis suggests that small RNA competition arises from limited enzymatic resources. Therefore, a model was created in order to gain insights into this competition.\n\nAbstractSmall RNAs have been determined to have an essential role in gene regulation. However, competition between small RNAs is a poorly understood aspect of small RNA dynamics. Recent evidence has suggested that competition between small RNA pathways arises from a scarcity of common resources essential for small RNA activity. In order to understand how competition affects small RNAs in C. elegans, a system of differential equations was used. The model recreates normal behavior of small RNAs and uses random sampling in order to determine the coefficients of competition for each small RNA class. The model includes endogenous small-interfering RNAs (endo-siRNA), exogenous small-interfering RNAs (exo-siRNA), and microRNAs (miRNA). The model predicts that exo-siRNAs is dominated by competition between endo-siRNAs and miRNAs. Furthermore, the model predicts that competition is required for normal levels of endogenous small RNAs to be maintained. Although the model makes several assumptions about cell dynamics, the model is still useful in order to understand competition between small RNA pathways.

Bioinformatics