Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Salmon provides accurate, fast, and bias-aware transcript expression estimates using dual-phase inference

We introduce Salmon, a new method for quantifying transcript abundance from RNA-seq reads that is highly-accurate and very fast. Salmon is the first transcriptome-wide quantifier to model and correct for fragment GC content bias, which we demonstrate substantially improves the accuracy of abundance estimates and the reliability of subsequent differential expression analysis compared to existing methods that do not account for these biases. Salmon achieves its speed and accuracy by combining a new dual-phase parallel inference algorithm and feature-rich bias models with an ultra-fast read mapping procedure. These innovations yield both exceptional accuracy and order-of-magnitude speed benefits over alignment-based methods.

Bioinformatics

TransRate: reference free quality assessment of de-novo transcriptome assemblies

TransRate is a tool for reference-free quality assessment of de novo transcriptome assemblies. Using only sequenced reads as the input, TransRate measures the quality of individual contigs and whole assemblies, enabling assembly optimization and comparison. TransRate can accurately evaluate assemblies of conserved and novel RNA molecules of any kind in any species. We show that it is more accurate than comparable methods and demonstrate its use on a variety of data.

Bioinformatics

Chromatin interactions correlate with local transcriptional activity in Saccharomyces cerevisiae

Genome organization is crucial for efficiently responding to DNA damage and regulating transcription. In this study, we relate the genome organization of Saccharomyces cerevisiae (budding yeast) to its transcription activity by analyzing published circularized chromosome conformation capture (4C) data in conjunction with eight separate datasets describing genome-wide transcription rate or RNA polymerase II (Pol II) occupancy. We find that large chromosome segments are more likely to interact in areas that have high transcription rate or Pol II occupancy. Additionally, we find that groups of genes with similar transcription rates or similar Pol II occupancy are more likely to have higher numbers of chromosomal interactions than groups of random genes. We hypothesize that transcription localization occurs around sets of genes with similar transcription rates, and more often around genes that are highly transcribed, in order to produce more efficient transcription. Our analysis cannot discern whether gene co-localization occurs because of similar transcription rates or whether similar transcription rates are a consequence of co-localization.

Bioinformatics

Clustering of mRNA-Seq data for detection of alternative splicing patterns

Current sequencing of mRNA can provide estimates of the levels of individual isoforms within the cell, where isoforms are the different distinct mRNA products or proteins created by a gene. It remains to adapt many standard statistical methods commonly used for analyzing gene expression levels to take advantage of this additional information. One novel question is whether we can find groupings or clusters of samples that are distinguished not by their gene expression but by their isoform usage. Such clusters in tumors, for example, could be the result of shared disruption to the splicing system that creates the different isoforms. We propose a novel approach to clustering mRNA-Seq data that identifies clusters of samples with common isoform usage. We show via simulation that our methods are more sensitive to finding clusters of similar alternative splicing patterns than standard clustering techniques applied directly to the estimates of isoform levels. We further demonstrate that clustering on isoform usage is more accurate than clustering directly on isoform levels by examining real data that contains a technical artifact that resulted in different batches having different isoform usage patterns. Clustering, mRNA-Seq, Alternative splicing

Bioinformatics

SARTools: a DESeq2- and edgeR-based R pipeline for comprehensive differential analysis of RNA-Seq data

BackgroundSeveral R packages exist for the detection of differentially expressed genes from RNA-Seq data. The analysis process includes three main steps, namely normalization, dispersion estimation and test for differential expression. Quality control steps along this process are recommended but not mandatory, and failing to check the characteristics of the dataset may lead to spurious results. In addition, normalization methods and statistical models are not exchangeable across the packages without adequate transformations the users are often not aware of. Thus, dedicated analysis pipelines are needed to include systematic quality control steps and prevent errors from misusing the proposed methods.\n\nResultsSARTools is an R pipeline for differential analysis of RNA-Seq count data. It can handle designs involving two or more conditions of a single biological factor with or without a blocking factor (such as a batch effect or a sample pairing). It is based on DESeq2 and edgeR and is composed of an R package and two R script templates (for DESeq2 and edgeR respectively). Tuning a small number of parameters and executing one of the R scripts, users have access to the full results of the analysis, including lists of differentially expressed genes and a HTML report that (i) displays diagnostic plots for quality control and model hypotheses checking and (ii) keeps track of the whole analysis process, parameter values and versions of the R packages used.\n\nConclusionsSARTools provides systematic quality controls of the dataset as well as diagnostic plots that help to tune the model parameters. It gives access to the main parameters of DESeq2 and edgeR and prevents untrained users from misusing some functionalities of both packages. By keeping track of all the parameters of the analysis process it fits the requirements of reproducible research.

Bioinformatics

Quantitative Relations in Protein and RNA Folding Deduced from Quantum Theory

Quantitative relations in protein and RNA folding are deduced from the quantum folding theory of macromolecules. It includes: deduction of the law on the temperature-dependence of folding rate and its tests on protein dataset; study on the chain-length dependence of the folding rate for a large class of biomolecules; deduction of the statistical relation of folding free energy versus chain-length; and deduction of the statistical relation between folding rate and chain length and its test on protein and RNA dataset. In the above quantum approach the influence of the solvent environment factor on folding rate has been taken into account automatically. The successes of the deduction of these new relations from the first principle and their successful comparison with experimental data afford strong evidence on the possible existence of a common quantum mechanism in the conformational change of biomolecules.

Bioinformatics

On the identifiability of transmission dynamic models for infectious diseases

Understanding the transmission dynamics of infectious diseases is important for both biological research and public health applications. It has been widely demonstrated that statistical modeling provides a firm basis for inferring relevant epidemiological quantities from incidence and molecular data. However, the complexity of transmission dynamic models causes two challenges: Firstly, the likelihood function of the models is generally not computable and computationally intensive simulation-based inference methods need to be employed. Secondly, the model may not be fully identifiable from the available data. While the first difficulty can be tackled by computational and algorithmic advances, the second obstacle is more fundamental. Identifiability issues may lead to inferences which are more driven by the prior assumptions than the data themselves. We here consider a popular and relatively simple, yet analytically intractable model for the spread of tuberculosis based on classical IS6110 fingerprinting data. We report on the identifiability of the model, presenting also some methodological advances regarding the inference. Using likelihood approximations, it is shown that the reproductive value cannot be identified from the data available and that the posterior distributions obtained in previous work have likely been substantially dominated by the assumed prior distribution. Further, we show that the inferences are influenced by the assumed infectious population size which has generally been kept fixed in previous work. We demonstrate that the infectious population size can be inferred if the remaining epidemiological parameters are already known with sufficient precision.

Bioinformatics

Livestock market data for modeling disease spread among US cattle

Transportation of livestock carries the risk of spreading foreign animal diseases, leading to costly public and private sector expenditures on disease containment and eradication. Livestock movement tracing systems in Europe, Australia and Japan have allowed epidemiologists to model the risks engendered by transportation of live animals and prepare responses designed to protect the livestock industry. Within the US, data on livestock movement is not sufficient for direct parameterization of models for disease spread, but network models that assimilate limited data provide a path forward in model development to inform preparedness for disease outbreaks in the US. Here, we develop a novel data stream, the information publicly reported by US livestock markets on the origin of cattle consigned at live auctions, and demonstrate the potential for estimating a national-scale network model of cattle movement. By aggregating auction reports generated weekly at markets in several states, including some archived reports spanning several years, we obtain a market-oriented sample of edges from the dynamic cattle transportation network in the US. We first propose a sampling framework that allows inference about shipments originating from operations not explicitly sampled and consigned at non-reporting livestock markets in the US, and we report key predictors that are influential in extrapolating beyond our opportunistic sample. As a demonstration of the utility gained from the data and fitted parameters, we model the critical role of market biosecurity procedures in the context of a spatially homogeneous but temporally dynamic representation of cattle movements following an introduction of a foreign animal disease. We conclude that auction market data fills critical gaps in our ability to model intrastate cattle movement for infectious disease dynamics, particularly with an ability to addresses the capacity of markets to amplify or control a livestock disease outbreak.\n\nAuthor SummaryWe have automated the collection of previously unavailable cattle movement data, allowing us to aggregate details on the origins of cattle sold at live-auction markets in the US. Using our novel dataset, we demonstrate potential to infer a complete dynamic transportation network that would drive disease transmission in models of potential US livestock epidemics.

Bioinformatics

Learning from heterogeneous data sources: an application in spatial proteomics

Sub-cellular localisation of proteins is an essential post-translational regulatory mechanism that can be assayed using high-throughput mass spectrometry (MS). These MS-based spatial proteomics experiments enable us to pinpoint the sub-cellular distribution of thousands of proteins in a specific system under controlled conditions. Recent advances in high-throughput MS methods have yielded a plethora of experimental spatial proteomics data for the cell biology community. Yet, there are many third-party data sources, such as immunofluorescence microscopy or protein annotations and sequences, which represent a rich and vast source of complementary information. We present a unique transfer learning classification framework that utilises a nearest-neighbour or support vector machine system, to integrate heterogeneous data sources to considerably improve on the quantity and quality of sub-cellular protein assignment. We demonstrate the utility of our algorithms through evaluation of five experimental datasets, from four different species in conjunction with four different auxiliary data sources to classify proteins to tens of sub-cellular compartments with high generalisation accuracy. We further apply the method to an experiment on pluripotent mouse embryonic stem cells to classify a set of previously unknown proteins, and validate our findings against a recent high resolution map of the mouse stem cell proteome. The methodology is distributed as part of the open-source Bioconductor pRoloc suite for spatial proteomics data analysis.\n\nAbbreviations

Bioinformatics

HTS-IBIS: fast and accurate inference of binding site motifs from HT-SELEX data

SummaryRecent technological advancements enable measuring the binding of a transcription factor to thousands of DNA sequences, in order to infer its binding preferences. High-throughput-SELEX measures protein-DNA binding by deep sequencing over several cycles of enrichment. We devised a new algorithm called HTS-IBIS for the inference task. HTS-IBIS corrects for technological biases, selects the cycle and k, and builds a motif starting from a consensus k-mer in that cycle. In large scale tests, HTS-IBIS outperformed the extant automatic algorithm for the motif finding task on both in vitro and in vivo binding prediction.\n\nAvailabilityHTS-IBIS is available on acgt.cs.tau.ac.il/HTS-IBIS.\n\nContactrshamir@tau.ac.il

Bioinformatics

Protein binding and methylation on looping chromatin accurately predict distal regulatory interactions

Identifying the gene targets of distal regulatory sequences is a challenging problem with the potential to illuminate the causal underpinnings of complex diseases. However, current experimental methods to map enhancer-promoter interactions genome-wide are limited by their cost and complexity. We present TargetFinder, a computational method that reconstructs a cells three-dimensional regulatory landscape from two-dimensional genomic features. TargetFinder achieves outstanding predictive accuracy across diverse cell lines with a false discovery rate up to fifteen times smaller than common heuristics, and reveals that distal regulatory interactions are characterized by distinct signatures of protein interactions and epigenetic marks on the DNA loop between an active enhancer and targeted promoter. Much of this signature is shared across cell types, shedding light on the role of chromatin organization in gene regulation and establishing TargetFinder as a method to accurately map long-range regulatory interactions using a small number of easily acquired datasets.

Bioinformatics

Fast and efficient QTL mapper for thousands of molecular phenotypes

MotivationIn order to discover quantitative trait loci (QTLs), multi-dimensional genomic data sets combining DNA-seq and ChiP-/RNA-seq require methods that rapidly correlate tens of thousands of molecular phenotypes with millions of genetic variants while appropriately controlling for multiple testing.\n\nResultsWe have developed FastQTL, a method that implements a popular cis-QTL mapping strategy in a user- and cluster-friendly tool. FastQTL also proposes an efficient permutation procedure to control for multiple testing. The outcome of permutations is modeled using beta distributions trained from a few permutations and from which adjusted p-values can be estimated at any level of significance with little computational cost. The Geuvadis & GTEx pilot data sets can be now easily analyzed an order of magnitude faster than previous approaches.\n\nAvailabilitySource code, binaries and comprehensive documentation of FastQTL are freely available to download at http://fastqtl.sourceforge.net/.\n\nContactolivier.delaneau@unige.ch

Bioinformatics

EC-PSI: Associating Enzyme Commission Numbers with Pfam Domains

With the growing number of protein structures in the protein data bank (PDB), there is a need to annotate these structures at the domain level in order to relate protein structure to protein function. Thanks to the SIFTS database, many PDB chains are now cross-referenced with Pfam domains and enzyme commission (EC) numbers. However, these annotations do not include any explicit relationship between individual Pfam domains and EC numbers. This article presents a novel statistical training-based method called EC-PSI that can automatically infer high confidence associations between EC numbers and Pfam domains directly from EC-chain associations from SIFTS and from EC-sequence associations from the SwissProt, and TrEMBL databases. By collecting and integrating these existing EC-chain/sequence annotations, our approach is able to infer a total of 8,329 direct EC-Pfam associations with an overall F-measure of 0.819 with respect to the manually curated InterPro database, which we treat here as a \"gold standard\" reference dataset. Thus, compared to the 1,493 EC-Pfam associations in InterPro, our approach provides a way to find over six times as many high quality EC-Pfam associations completely automatically.

Bioinformatics

A max-margin model for predicting residue-base contacts in protein-RNA interactions

Protein-RNA interactions (PRIs) are essential for many biological processes, so understanding aspects of the sequences and structures involved in PRIs is important for unraveling such processes. Because of the expensive and time-consuming techniques required for experimental determination of complex protein-RNA structures, various computational methods have been developed to predict PRIs. However, most of these methods focus on predicting only RNA-binding regions in proteins or only protein-binding motifs in RNA. Methods for predicting entire residue-base contacts in PRIs have not yet achieved sufficient accuracy. Furthermore, some of these methods require the identification of 3D structures or homologous sequences, which are not available for all protein and RNA sequences. Here, we propose a prediction method for predicting residue-base contacts between proteins and RNAs using only sequence information and structural information predicted from sequences. The method can be applied to any protein-RNA pair, even when rich information such as its 3D structure, is not available. In this method, residue-base contact prediction is formalized as an integer programming problem. We predict a residue-base contact map that maximizes a scoring function based on sequence-based features such as k-mers of sequences and the predicted secondary structure. The scoring function is trained using a max-margin framework from known PRIs with 3D structures. To verify our method, we conducted several computational experiments. The results suggest that our method, which is based on only sequence information, is comparable with RNA-binding residue prediction methods based on known binding data.

Bioinformatics

A profile-based method for identifying functional divergence of orthologous genes in bacterial genomes

MotivationNext generation sequencing technologies have provided us with a wealth of information on genetic variation, but predicting the functional significance of this variation is a difficult task. While many comparative genomics studies have focused on gene flux and large scale changes, relatively little attention has been paid to quantifying the effects of single nucleotide polymorphisms and indels on protein function, particularly in bacterial genomics.\n\nResultsWe present a hidden Markov model based approach we call delta-bitscore (DBS) for identifying orthologous proteins that have diverged at the amino acid sequence level in a way that is likely to impact biological function. We benchmark this approach with several widely used datasets and apply it to a proof-of-concept study of orthologous proteomes in an investigation of host adaptation in Salmonella enterica. We highlight the value of the method in identifying functional divergence of genes, and suggest that this tool may be a better approach than the commonly used dN/dS metric for identifying functionally significant genetic changes occurring in recently diverged organisms.\n\nAvailabilityA program implementing DBS for pairwise genome comparisons is freely available at: https://github.com/UCanCompBio/deltaBS.\n\nContactnicole.wheeler@pg.canterbury.ac.nz, lars.barquist@uni-wuerzburg.de\n\nSupplementary informationSupplementary data are available at BioRxiv online.

Bioinformatics

metaCCA: Summary statistics-based multivariate meta-analysis of genome-wide association studies using canonical correlation analysis

A dominant approach to genetic association studies is to perform univariate tests between genotype-phenotype pairs. However, analysing related traits together increases statistical power, and certain complex associations become detectable only when several variants are tested jointly. Currently, modest sample sizes of individual cohorts and restricted availability of individual-level genotype-phenotype data across the cohorts limit conducting multivariate tests.\n\nWe introduce metaCCA, a computational framework for summary statistics-based analysis of a single or multiple studies that allows multivariate representation of both genotype and phenotype. It extends the statistical technique of canonical correlation analysis to the setting where original individual-level records are not available, and employs a covariance shrinkage algorithm to achieve robustness.\n\nMultivariate meta-analysis of two Finnish studies of nuclear magnetic resonance metabolomics by metaCCA, using standard univariate output from the program SNPTEST, shows an excellent agreement with the pooled individual-level analysis of original data. Motivated by strong multivariate signals in the lipid genes tested, we envision that multivariate association testing using metaCCA has a great potential to provide novel insights from already published summary statistics from high-throughput phenotyping technologies.\n\nCode is available at https://github.com/aalto-ics-kepaco.

Bioinformatics

Tools and techniques for computational reproducibility

When reporting research findings, scientists document the steps they followed so that others can verify and build upon the research. When those steps have been described in sufficient detail that others can retrace the steps and obtain similar results, the research is said to be reproducible. Computers play a vital role in many research disciplines and present both opportunities and challenges for reproducibility. Computers can be programmed to execute analysis tasks, and those programs can be repeated and shared with others. Due to the deterministic nature of most computer programs, the same analysis tasks, applied to the same data, will often produce the same outputs. However, in practice, computational findings often cannot be reproduced due to complexities in how software is packaged, installed, and executed--and due to limitations in how scientists document analysis steps. Many tools and techniques are available to help overcome these challenges. Here we describe seven such strategies. With a broad scientific audience in mind, we describe strengths and limitations of each approach, as well as circumstances under which each might be applied. No single strategy is sufficient for every scenario; thus we emphasize that it is often useful to combine approaches.

Bioinformatics

Scalable multi whole-genome alignment using recursive exact matching

The emergence of third generation sequencing technologies has brought near perfect de-novo genome assembly within reach. This clears the way towards reference-free detection of genomic variations.\n\nIn this paper, we introduce a novel concept for aligning whole-genomes which allows the alignment of multiple genomes. Alignments are constructed in a recursive manner, in which alignment decisions are statistically supported. Computational performance is achieved by splitting an initial indexing data structure into a multitude of smaller indices.\n\nWe show that our method can be used to detect high resolution structural variations between two human genomes, and that it can be used to obtain a high quality multiple genome alignment of at least nineteen Mycobacterium tuberculosis genomes.\n\nAn implementation of the outlined algorithm called REVEAL is available on: https://github.com/jasperlinthorst/REVEAL

Bioinformatics