Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Discovery of methylation loci and analyses of differential methylation from replicated high-throughput sequencing data

Cytosine methylation is widespread in most eukaryotic genomes and is known to play a substantial role in various regulatory pathways. Unmethylated cytosines may be converted to uracil through the addition of sodium bisulphite, allowing genome-wide quantification of cytosine methylation via high-throughput sequencing. The data thus acquired allows the discovery of methylation loci; contiguous regions of methylation consistently methylated across biological replicates. The mapping of these loci allows for associations with other genomic factors to be identified, and for analyses of differential methylation to take place.\n\nThe segmentSeq R package is extended to identify methylation loci from high-throughput sequencing data from multiple experimental conditions. A statistical model is then developed that accounts for biological replication and variable rates of non-conversion of cytosines in each sample to compute posterior likelihoods of methylation at each locus within an empirical Bayesian framework. The same model is used as a basis for analysis of differential methylation between multiple experimental conditions with the baySeq R package. We demonstrate this method through an analysis of data derived from Dicer-like mutants in Arabidopsis that reveals complex interactions between the different Dicer-like mutants and their methylation pathways. We also show in simulation studies that this approach can be significantly more powerful in the detection of differential methylation than existing methods.

Bioinformatics

On enhancing variation detection through pan-genome indexing

Detection of genomic variants is commonly conducted by aligning a set of reads sequenced from an individual to the reference genome of the species and analyzing the resulting read pileup. Typically, this process finds a subset of variants already reported in databases and additional novel variants characteristic to the sequenced individual. Most of the effort in the literature has been put to the alignment problem on a single reference sequence, although our gathered knowledge on species such as human is pan-genomic: We know most of the common variation in addition to the reference sequence. There have been some efforts to exploit pan-genome indexing, where the most widely adopted approach is to build an index structure on a set of reference sequences containing observed variation combinations.\n\nThe enhancement in alignment accuracy when using pan-genome indexing has been demonstrated in experiments, but so far the above multiple references pan-genome indexing approach has not been tested on its final goal, that is, in enhancing variation detection. This is the focus of this article: We study a generic approach to add variation detection support on top of the multiple references pan-genomic indexing approach. Namely, we study the read pileup on a multiple alignment of reference genomes, and propose a heaviest path algorithm to extract a new recombined reference sequence. This recombined reference sequence can then be utilized in any standard read alignment and variation detection workflow. We demonstrate that the approach enhances variation detection on realistic data sets.

Bioinformatics

SSCM: A method to analyze and predict the pathogenicity of sequence variants

AbstractThe advent of cost-effective DNA sequencing has provided clinics with high-resolution information about patients genetic variants, which has resulted in the need for efficient interpretation of this genomic data. Traditionally, variant interpretation has been dominated by many manual, time-consuming processes due to the disparate forms of relevant information in clinical databases and literature. Computational techniques promise to automate much of this, and while they currently play only a supporting role, their continued improvement for variant interpretation is necessary to tackle the problem of scaling genetic sequencing to ever larger populations. Here, we present SSCM-Pathogenic, a genome-wide, allele-specific score for predicting variant pathogenicity. The score, generated by a semi-supervised clustering algorithm, shows predictive power on clinically relevant mutations, while also displaying predictive ability in noncoding regions of the genome.

Bioinformatics

In Silico Predictive Modeling of CRISPR/Cas9 guide efficiency

The CRISPR/Cas9 system provides unprecedented genome editing capabilities; however, several facets of this system are under investigation for further characterization and optimization, including the choice of guide RNA that directs Cas9 to target DNA. In particular, given that one would like to target the protein-coding region of a gene, hundreds of guides satisfy the basic constraints of the CRISPR/Cas9 Protospacer Adjacent Motif sequence (PAM); however, not all of these guides actually generate gene knockouts with equal efficiency. Leveraging a broad set of experimental measurements of guide knockout efficiency, we introduce a state-of-the art in silico modeling approach to identify guides that will lead to more effective gene knockout. We first investigated which guide and gene features are critical for prediction (e.g., single- and di-nucleotide identity of the gene target), which are helpful (e.g., thermodynamics), and which are predictive but redundant (e.g., microhomology). We also investigated evaluation measures for comparing predictive models in the present context, suggesting that Area Under the Receiver Operating Curve is not ideal. Finally, we explored a variety of different model classes and found that use of gradient-boosted regression trees produced the best predictive performance. Pointers to our open-source software, code, and prediction server will be available at http://research.microsoft.com/en-us/projects/azimuth.

Bioinformatics

Modeling small RNA competition in C. elegans

Website SummarySmall RNAs are important regulators of gene expression; however, the relationship between small RNAs is poorly understood. Studying the crosstalk between small RNA pathways can help understand how gene expression is regulated. One hypothesis suggests that small RNA competition arises from limited enzymatic resources. Therefore, a model was created in order to gain insights into this competition.\n\nAbstractSmall RNAs have been determined to have an essential role in gene regulation. However, competition between small RNAs is a poorly understood aspect of small RNA dynamics. Recent evidence has suggested that competition between small RNA pathways arises from a scarcity of common resources essential for small RNA activity. In order to understand how competition affects small RNAs in C. elegans, a system of differential equations was used. The model recreates normal behavior of small RNAs and uses random sampling in order to determine the coefficients of competition for each small RNA class. The model includes endogenous small-interfering RNAs (endo-siRNA), exogenous small-interfering RNAs (exo-siRNA), and microRNAs (miRNA). The model predicts that exo-siRNAs is dominated by competition between endo-siRNAs and miRNAs. Furthermore, the model predicts that competition is required for normal levels of endogenous small RNAs to be maintained. Although the model makes several assumptions about cell dynamics, the model is still useful in order to understand competition between small RNA pathways.

Bioinformatics

Salmon provides accurate, fast, and bias-aware transcript expression estimates using dual-phase inference

We introduce Salmon, a new method for quantifying transcript abundance from RNA-seq reads that is highly-accurate and very fast. Salmon is the first transcriptome-wide quantifier to model and correct for fragment GC content bias, which we demonstrate substantially improves the accuracy of abundance estimates and the reliability of subsequent differential expression analysis compared to existing methods that do not account for these biases. Salmon achieves its speed and accuracy by combining a new dual-phase parallel inference algorithm and feature-rich bias models with an ultra-fast read mapping procedure. These innovations yield both exceptional accuracy and order-of-magnitude speed benefits over alignment-based methods.

Bioinformatics

TransRate: reference free quality assessment of de-novo transcriptome assemblies

TransRate is a tool for reference-free quality assessment of de novo transcriptome assemblies. Using only sequenced reads as the input, TransRate measures the quality of individual contigs and whole assemblies, enabling assembly optimization and comparison. TransRate can accurately evaluate assemblies of conserved and novel RNA molecules of any kind in any species. We show that it is more accurate than comparable methods and demonstrate its use on a variety of data.

Bioinformatics

Chromatin interactions correlate with local transcriptional activity in Saccharomyces cerevisiae

Genome organization is crucial for efficiently responding to DNA damage and regulating transcription. In this study, we relate the genome organization of Saccharomyces cerevisiae (budding yeast) to its transcription activity by analyzing published circularized chromosome conformation capture (4C) data in conjunction with eight separate datasets describing genome-wide transcription rate or RNA polymerase II (Pol II) occupancy. We find that large chromosome segments are more likely to interact in areas that have high transcription rate or Pol II occupancy. Additionally, we find that groups of genes with similar transcription rates or similar Pol II occupancy are more likely to have higher numbers of chromosomal interactions than groups of random genes. We hypothesize that transcription localization occurs around sets of genes with similar transcription rates, and more often around genes that are highly transcribed, in order to produce more efficient transcription. Our analysis cannot discern whether gene co-localization occurs because of similar transcription rates or whether similar transcription rates are a consequence of co-localization.

Bioinformatics

Clustering of mRNA-Seq data for detection of alternative splicing patterns

Current sequencing of mRNA can provide estimates of the levels of individual isoforms within the cell, where isoforms are the different distinct mRNA products or proteins created by a gene. It remains to adapt many standard statistical methods commonly used for analyzing gene expression levels to take advantage of this additional information. One novel question is whether we can find groupings or clusters of samples that are distinguished not by their gene expression but by their isoform usage. Such clusters in tumors, for example, could be the result of shared disruption to the splicing system that creates the different isoforms. We propose a novel approach to clustering mRNA-Seq data that identifies clusters of samples with common isoform usage. We show via simulation that our methods are more sensitive to finding clusters of similar alternative splicing patterns than standard clustering techniques applied directly to the estimates of isoform levels. We further demonstrate that clustering on isoform usage is more accurate than clustering directly on isoform levels by examining real data that contains a technical artifact that resulted in different batches having different isoform usage patterns. Clustering, mRNA-Seq, Alternative splicing

Bioinformatics

SARTools: a DESeq2- and edgeR-based R pipeline for comprehensive differential analysis of RNA-Seq data

BackgroundSeveral R packages exist for the detection of differentially expressed genes from RNA-Seq data. The analysis process includes three main steps, namely normalization, dispersion estimation and test for differential expression. Quality control steps along this process are recommended but not mandatory, and failing to check the characteristics of the dataset may lead to spurious results. In addition, normalization methods and statistical models are not exchangeable across the packages without adequate transformations the users are often not aware of. Thus, dedicated analysis pipelines are needed to include systematic quality control steps and prevent errors from misusing the proposed methods.\n\nResultsSARTools is an R pipeline for differential analysis of RNA-Seq count data. It can handle designs involving two or more conditions of a single biological factor with or without a blocking factor (such as a batch effect or a sample pairing). It is based on DESeq2 and edgeR and is composed of an R package and two R script templates (for DESeq2 and edgeR respectively). Tuning a small number of parameters and executing one of the R scripts, users have access to the full results of the analysis, including lists of differentially expressed genes and a HTML report that (i) displays diagnostic plots for quality control and model hypotheses checking and (ii) keeps track of the whole analysis process, parameter values and versions of the R packages used.\n\nConclusionsSARTools provides systematic quality controls of the dataset as well as diagnostic plots that help to tune the model parameters. It gives access to the main parameters of DESeq2 and edgeR and prevents untrained users from misusing some functionalities of both packages. By keeping track of all the parameters of the analysis process it fits the requirements of reproducible research.

Bioinformatics

Quantitative Relations in Protein and RNA Folding Deduced from Quantum Theory

Quantitative relations in protein and RNA folding are deduced from the quantum folding theory of macromolecules. It includes: deduction of the law on the temperature-dependence of folding rate and its tests on protein dataset; study on the chain-length dependence of the folding rate for a large class of biomolecules; deduction of the statistical relation of folding free energy versus chain-length; and deduction of the statistical relation between folding rate and chain length and its test on protein and RNA dataset. In the above quantum approach the influence of the solvent environment factor on folding rate has been taken into account automatically. The successes of the deduction of these new relations from the first principle and their successful comparison with experimental data afford strong evidence on the possible existence of a common quantum mechanism in the conformational change of biomolecules.

Bioinformatics

On the identifiability of transmission dynamic models for infectious diseases

Understanding the transmission dynamics of infectious diseases is important for both biological research and public health applications. It has been widely demonstrated that statistical modeling provides a firm basis for inferring relevant epidemiological quantities from incidence and molecular data. However, the complexity of transmission dynamic models causes two challenges: Firstly, the likelihood function of the models is generally not computable and computationally intensive simulation-based inference methods need to be employed. Secondly, the model may not be fully identifiable from the available data. While the first difficulty can be tackled by computational and algorithmic advances, the second obstacle is more fundamental. Identifiability issues may lead to inferences which are more driven by the prior assumptions than the data themselves. We here consider a popular and relatively simple, yet analytically intractable model for the spread of tuberculosis based on classical IS6110 fingerprinting data. We report on the identifiability of the model, presenting also some methodological advances regarding the inference. Using likelihood approximations, it is shown that the reproductive value cannot be identified from the data available and that the posterior distributions obtained in previous work have likely been substantially dominated by the assumed prior distribution. Further, we show that the inferences are influenced by the assumed infectious population size which has generally been kept fixed in previous work. We demonstrate that the infectious population size can be inferred if the remaining epidemiological parameters are already known with sufficient precision.

Bioinformatics

Livestock market data for modeling disease spread among US cattle

Transportation of livestock carries the risk of spreading foreign animal diseases, leading to costly public and private sector expenditures on disease containment and eradication. Livestock movement tracing systems in Europe, Australia and Japan have allowed epidemiologists to model the risks engendered by transportation of live animals and prepare responses designed to protect the livestock industry. Within the US, data on livestock movement is not sufficient for direct parameterization of models for disease spread, but network models that assimilate limited data provide a path forward in model development to inform preparedness for disease outbreaks in the US. Here, we develop a novel data stream, the information publicly reported by US livestock markets on the origin of cattle consigned at live auctions, and demonstrate the potential for estimating a national-scale network model of cattle movement. By aggregating auction reports generated weekly at markets in several states, including some archived reports spanning several years, we obtain a market-oriented sample of edges from the dynamic cattle transportation network in the US. We first propose a sampling framework that allows inference about shipments originating from operations not explicitly sampled and consigned at non-reporting livestock markets in the US, and we report key predictors that are influential in extrapolating beyond our opportunistic sample. As a demonstration of the utility gained from the data and fitted parameters, we model the critical role of market biosecurity procedures in the context of a spatially homogeneous but temporally dynamic representation of cattle movements following an introduction of a foreign animal disease. We conclude that auction market data fills critical gaps in our ability to model intrastate cattle movement for infectious disease dynamics, particularly with an ability to addresses the capacity of markets to amplify or control a livestock disease outbreak.\n\nAuthor SummaryWe have automated the collection of previously unavailable cattle movement data, allowing us to aggregate details on the origins of cattle sold at live-auction markets in the US. Using our novel dataset, we demonstrate potential to infer a complete dynamic transportation network that would drive disease transmission in models of potential US livestock epidemics.

Bioinformatics

Learning from heterogeneous data sources: an application in spatial proteomics

Sub-cellular localisation of proteins is an essential post-translational regulatory mechanism that can be assayed using high-throughput mass spectrometry (MS). These MS-based spatial proteomics experiments enable us to pinpoint the sub-cellular distribution of thousands of proteins in a specific system under controlled conditions. Recent advances in high-throughput MS methods have yielded a plethora of experimental spatial proteomics data for the cell biology community. Yet, there are many third-party data sources, such as immunofluorescence microscopy or protein annotations and sequences, which represent a rich and vast source of complementary information. We present a unique transfer learning classification framework that utilises a nearest-neighbour or support vector machine system, to integrate heterogeneous data sources to considerably improve on the quantity and quality of sub-cellular protein assignment. We demonstrate the utility of our algorithms through evaluation of five experimental datasets, from four different species in conjunction with four different auxiliary data sources to classify proteins to tens of sub-cellular compartments with high generalisation accuracy. We further apply the method to an experiment on pluripotent mouse embryonic stem cells to classify a set of previously unknown proteins, and validate our findings against a recent high resolution map of the mouse stem cell proteome. The methodology is distributed as part of the open-source Bioconductor pRoloc suite for spatial proteomics data analysis.\n\nAbbreviations

Bioinformatics

HTS-IBIS: fast and accurate inference of binding site motifs from HT-SELEX data

SummaryRecent technological advancements enable measuring the binding of a transcription factor to thousands of DNA sequences, in order to infer its binding preferences. High-throughput-SELEX measures protein-DNA binding by deep sequencing over several cycles of enrichment. We devised a new algorithm called HTS-IBIS for the inference task. HTS-IBIS corrects for technological biases, selects the cycle and k, and builds a motif starting from a consensus k-mer in that cycle. In large scale tests, HTS-IBIS outperformed the extant automatic algorithm for the motif finding task on both in vitro and in vivo binding prediction.\n\nAvailabilityHTS-IBIS is available on acgt.cs.tau.ac.il/HTS-IBIS.\n\nContactrshamir@tau.ac.il

Bioinformatics

Protein binding and methylation on looping chromatin accurately predict distal regulatory interactions

Identifying the gene targets of distal regulatory sequences is a challenging problem with the potential to illuminate the causal underpinnings of complex diseases. However, current experimental methods to map enhancer-promoter interactions genome-wide are limited by their cost and complexity. We present TargetFinder, a computational method that reconstructs a cells three-dimensional regulatory landscape from two-dimensional genomic features. TargetFinder achieves outstanding predictive accuracy across diverse cell lines with a false discovery rate up to fifteen times smaller than common heuristics, and reveals that distal regulatory interactions are characterized by distinct signatures of protein interactions and epigenetic marks on the DNA loop between an active enhancer and targeted promoter. Much of this signature is shared across cell types, shedding light on the role of chromatin organization in gene regulation and establishing TargetFinder as a method to accurately map long-range regulatory interactions using a small number of easily acquired datasets.

Bioinformatics

Fast and efficient QTL mapper for thousands of molecular phenotypes

MotivationIn order to discover quantitative trait loci (QTLs), multi-dimensional genomic data sets combining DNA-seq and ChiP-/RNA-seq require methods that rapidly correlate tens of thousands of molecular phenotypes with millions of genetic variants while appropriately controlling for multiple testing.\n\nResultsWe have developed FastQTL, a method that implements a popular cis-QTL mapping strategy in a user- and cluster-friendly tool. FastQTL also proposes an efficient permutation procedure to control for multiple testing. The outcome of permutations is modeled using beta distributions trained from a few permutations and from which adjusted p-values can be estimated at any level of significance with little computational cost. The Geuvadis & GTEx pilot data sets can be now easily analyzed an order of magnitude faster than previous approaches.\n\nAvailabilitySource code, binaries and comprehensive documentation of FastQTL are freely available to download at http://fastqtl.sourceforge.net/.\n\nContactolivier.delaneau@unige.ch

Bioinformatics

EC-PSI: Associating Enzyme Commission Numbers with Pfam Domains

With the growing number of protein structures in the protein data bank (PDB), there is a need to annotate these structures at the domain level in order to relate protein structure to protein function. Thanks to the SIFTS database, many PDB chains are now cross-referenced with Pfam domains and enzyme commission (EC) numbers. However, these annotations do not include any explicit relationship between individual Pfam domains and EC numbers. This article presents a novel statistical training-based method called EC-PSI that can automatically infer high confidence associations between EC numbers and Pfam domains directly from EC-chain associations from SIFTS and from EC-sequence associations from the SwissProt, and TrEMBL databases. By collecting and integrating these existing EC-chain/sequence annotations, our approach is able to infer a total of 8,329 direct EC-Pfam associations with an overall F-measure of 0.819 with respect to the manually curated InterPro database, which we treat here as a \"gold standard\" reference dataset. Thus, compared to the 1,493 EC-Pfam associations in InterPro, our approach provides a way to find over six times as many high quality EC-Pfam associations completely automatically.

Bioinformatics