Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

BRANE Cut: Biologically-Related A priori Network Enhancement with Graph cuts for Gene Regulatory Network Inference

Background: Inferring gene networks from high-throughput data constitutes an important step in the discovery of relevant regulatory relationships in organism cells. Despite the large number of available Gene Regulatory Network inference methods, the problem remains challenging: the underdetermination in the space of possible solutions requires additional constraints that incorporate a priori information on gene interactions. Methods: Weighting all possible pairwise gene relationships by a probability of edge presence, we formulate the regulatory network inference as a discrete variational problem on graphs. We enforce biologically plausible coupling between groups and types of genes by minimizing an edge labeling functional coding for a priori structures. The optimization is carried out with Graph cuts, an approach popular in image processing and computer vision. We compare the inferred regulatory networks to results achieved by the mutual-information-based Context Likelihood of Relatedness (CLR) method and by the state-of-the-art GENIE3, winner of the DREAM4 multifactorial challenge. Results: Our BRANE Cut approach infers more accurately the five DREAM4 in silico networks (with improvements from 6% to 11%). On a real Escherichia coli compendium, an improvement of 11.8% compared to CLR and 3% compared to GENIE3 is obtained in terms of Area Under Precision-Recall curve. Up to 48 additional verified interactions are obtained over GENIE3 for a given precision. On this dataset involving 4345 genes, our method achieves a performance similar to that of GENIE3, while being more than seven times faster. The BRANE Cut code is available at: http://www-syscom.univ-mlv.fr/~pirayre/Codes-GRN-BRANE-cut.html Conclusions: BRANE Cut is a weighted graph thresholding method. Using biologically sound penalties and data-driven parameters, it improves three state-of-the-art GRN inference methods. It is applicable as a generic network inference post-processing, due its computational efficiency.

Bioinformatics

ASAP: A Machine-Learning Framework for Local Protein Properties

Determining residue level protein properties, such as the sites for post-translational modifications (PTMs) are vital to understanding proteins at all levels of function. Experimental methods are costly and time-consuming, thus high confidence predictions become essential for functional knowledge at a genomic scale. Traditional computational methods based on strict rules (e.g. regular expressions) fail to annotate sites that lack substantial similarity. Thus, Machine Learning (ML) methods become fundamental in annotating proteins with unknown function. We present ASAP (Amino-acid Sequence Annotation Prediction), a universal ML framework for residue-level predictions. ASAP extracts efficiently and fast large set of window-based features from raw sequences. The platform also supports easy integration of external features such as secondary structure or PSSM profiles. The features are then combined to train underlying ML classifiers. We present a detailed case study for ASAP that was used to train CleavePred, a state-of-the-art protein precursor cleavage sites predictor. Protein cleavage is a fundamental PTM shared by a wide variety of protein groups with minimal sequence similarity. Current computational methods have high false positive rates, making them suboptimal for this task. CleavePred has a simple Python API, and is freely accessible via a web-based application. The high performance of ASAP toward the task of precursor cleavage is suited for analyzing new proteomes at a genomic scale. The tool is attractive to protein design, mass spectrometry search engines and the discovery of new peptide hormones. In summary, we illustrate ASAP as an entry point for predicting PTMs. The approach and flexibility of the platform can easily be extended for additional residue specific tasks. ASAP and CleavePred source code available at https://github.com/ddofer/asap.

Bioinformatics

BioJS-HGV Viewer: Genetic Variation Visualizer

MotivationStudying the pattern of genetic variants is a primary step in deciphering the basis of biological diversity, identifying key driver variants that affect disease states and evolution of a species. Catalogs of genetic variants contain vast numbers of variants and are growing at an exponential rate, but lack an interactive exploratory interface.\n\nResultsWe present BioJS-HGV Viewer, a BioJS component to represent and visualize genetic variants pooled from different sources. The tool displays sequences and variants at different levels of detail, facilitating representation of variant sites and annotations in a user friendly and interactive manner.\n\nAvailabilitySource code for BioJS-HGV Viewer is available at: https://github.com/saketkc/biojs-genetic-variation-viewer\n\nA demo is available at: http://saketkc.github.io/biojs-genetic-variation-viewer\n\nContact: martin@ebi.ac.uk

Bioinformatics

G-quadruplex prediction in E.coli genome reveals a conserved putative G-quadruplex-Hairpin-Duplex switch

Many studies show that short non-coding sequences are widely conserved among regulatory elements. More and more conserved sequences are being discovered since the development next generation sequencing technology. A common approach to identify conserved sequences with regulatory roles rely on the topological change such as hairpin formation on the DNA or RNA level. However, G-quadruplexes, a non-cannonical nucleic acid topology with little established biological role, is rarely considered for conserved regulatory element discovery. Here we present the use of G-quadruplex prediction algorithm to identify putative G-quadruplex-forming and conserved elements in E.coli genome. Phylogenetic analysis of 52 G-quadruplex forming sequences revealed two conserved G-quadruplex motifs with potential regulatory role.

Bioinformatics

Quantitative proteome-based guidelines for intrinsic disorder characterization

Intrinsically disordered proteins fail to adopt a stable three-dimensional structure under physiological conditions. It is now understood that many disordered proteins are not dysfunctional, but instead engage in numerous cellular processes, including signaling and regulation. Disorder characterization from amino acid sequence relies on computational disorder prediction algorithms. While numerous large-scale investigations of disorder have been performed using these algorithms, and have offered valuable insight regarding the prevalence of protein disorder in many organisms, critical proteome-based descriptive statistical guidelines that would enable the objective assessment of intrinsic disorder in a protein of interest remain to be established. Here we present a quantitative characterization of numerous disorder features using a rigorous non-parametric statistical approach, providing expected values and percentile cutoffs for each feature in ten eukaryotic proteomes. Our estimates utilize multiple ab initio disorder prediction algorithms grounded on physicochemical principles. Furthermore, we present novel threshold values, specific to both the prediction algorithms and the proteomes, defining the longest primary sequence length in which the significance of a continuous disordered region can be evaluated on the basis of length alone. The guidelines presented here are intended to improve the interpretation of disorder content and continuous disorder predictions from the proteomic point of view.

Bioinformatics

Quantifying polygenic effects in genome-wide association studies using generalized estimating equations

1Recently, LD Score regression1 has been proposed as a computationally fast method to contrast confounding biases with polygenicity and to quantify their contribution to the inflation of test statistics in GWAS.\n\nIn this communication, we extend the LD Score regression approach by applying the generalized estimation equations (GEE) framework, which is capable of incorporating more external information from reference panels about the correlation structure of test statistics. We apply our GEE approach and LD Score regression to simulated and real data to compare their performance.\n\nWe show that our proposed methodology obtains more efficient estimates while preserving the robustness and desired properties of LD Score regression.

Bioinformatics

How reliable are ligand-centric methods for Target Fishing?

Computational methods for Target Fishing (TF), also known as Target Prediction or Polypharmacology Prediction, can be used to discover new targets in small-molecule drugs. This may result in repositioning the drug in a new indication or improving our current understanding of its efficacy and side effects. While there is a substantial body of research on TF methods, there is still a need to improve their validation, which is often limited to a small part of the available targets and not easily interpretable by the user. Here we discuss how target-centric TF methods are inherently limited by the number of targets that can possibly predict (this number is by construction much larger in ligand-centric techniques). We also propose a new benchmark to validate TF methods, which is particularly suited to analyse how predictive performance varies with the query molecule. On average over approved drugs, we estimate that only five predicted targets will have to be tested to find two true targets with submicromolar potency (a strong variability in performance is however observed). In addition, we find that an approved drug has currently an average of eight known targets, which reinforces the notion that polypharmacology is a common and strong event. Furthermore, with the assistance of a control group of randomly-selected molecules, we show that the targets of approved drugs are generally harder to predict.

Bioinformatics

Rich chromatin structure prediction from Hi-C data

Recent studies involving the 3-dimensional conformation of chromatin have revealed the important role it has to play in different processes within the cell. These studies have also led to the discovery of densely interacting segments of the chromosome, called topologically associating domains. The accurate identification of these domains from Hi-C interaction data is an interesting and important computational problem for which numerous methods have been proposed. Unfortunately, most existing algorithms designed to identify these domains assume that they are non-overlapping whereas there is substantial evidence to believe a nested structure exists. We present an efficient methodology to predict hierarchical chromatin domains using chromatin conformation capture data. Our method predicts domains at different resolutions and uses these to construct a hierarchy that is based on intrinsic properties of the chromatin data. The hierarchy consists of a set of non-overlapping domains, that maximize intra-domain interaction frequencies, at each level. We show that our predicted structure is highly enriched for CTCF and various other chromatin markers. We also show that large-scale domains, at multiple resolutions within our hierarchy, are conserved across cell types and species. Our software, Matryoshka, is written in C++11 and licensed under GPL v3; it is available at https://github.com/COMBINE-lab/matryoshka.

Bioinformatics

Unearthing new genomic markers of drug response by improved measurement of discriminative power

BackgroundOncology drugs are only effective in a small proportion of cancer patients. Our current ability to identify these responsive patients before treatment is still poor in most cases. Thus, there is a pressing need to discover response markers for marketed and research oncology drugs in order to improve patient survival, reduce healthcare costs and enhance success rates in clinical trials. Screening these drugs against a large panel of cancer cell lines has been employed to discover new genomic markers of in vitro drug response, which can now be further evaluated on more accurate tumour models. However, while the identification of discriminative markers among thousands of candidate drug-gene associations in the data is error-prone, an appraisal of the effectiveness of such detection task is currently lacking.\n\nResultsHere we present a new non-parametric method to measuring the discriminative power of a drug-gene association. This is enabled by the identification of an auxiliary threshold posing this task as a binary classification problem. Unlike parametric statistical tests, the adopted non-parametric test has the advantage of not making strong assumptions about the data distorting the identification of genomic markers. Furthermore, we introduce a new benchmark to further validate these markers in vitro using more recent data not used to identify the markers. The application of this new methodology has led to the identification of 128 new genomic markers distributed across 61% of the analysed drugs, including 5 drugs without previously known markers, which were missed by the MANOVA test initially applied to analyse data from the Genomics of Drug Sensitivity in Cancer consortium.\n\nAbbreviation

Bioinformatics

Sequenceserver: a modern graphical user interface for custom BLAST databases

The dramatic drop in DNA sequencing costs has created many opportunities for novel biological research. These opportunities largely rest upon the ability to effectively compare newly obtained and previously known sequences. This is commonly done with BLAST, yet using BLAST directly on new datasets requires substantial technical skills or helpful colleagues. Furthermore, graphical interfaces for BLAST are challenging to install and largely mimic underlying computational processes rather than work patterns of researchers.\n\nWe combined a user-centric design philosophy with sustainable software development approaches to create Sequenceserver (http://sequenceserver.com), a modern graphical user interface for BLAST. Sequenceserver substantially increases the efficiency of researchers working with sequence data. This is due to innovations at three levels. First, our software can be installed and used on custom datasets extremely rapidly for personal and shared applications. Second, based on analysis of user input and simple algorithms, Sequenceserver reduces the amount of decisions the user must make, provides interactive visual feedback, and prevents common potential errors that would otherwise cause erroneous results. Finally, Sequenceserver provides multiple highly visual and text-based output options that mirror the requirements and work patterns of researchers. Together, these features greatly facilitate BLAST analysis and interpretation and thus substantially enhance researcher productivity.

Bioinformatics

Two Independent and Highly Efficient Open Source TKF91 Implementations

In the context of a master level programming practical at the computer science department of the Karlsruhe Institute of Technology, we developed and make available two independent and highly optimized open-source implementations for the pair-wise statistical alignment model, also known as TKF91, that was developed by Thorne, Kishino, and Felsenstein in 1991. This paper has two parts. In the educational part, we cover teaching issues regarding the setup of the course and the practical and summarize student and teacher experiences. In the scientific part, the two student teams (Team I: Nikolai, Sebastian, Daniel; Team II: Sarah, Pierre) present their solutions for implementing efficient and numerically stable implementations of the TKF91 algorithm. The two teams worked independently on implementing the same algorithm. Hence, since the implementations yield identical results --with slight numerical deviations-- we are confident that the implementations are correct. We describe the optimizations applied and make them available as open-source codes in the hope that our findings and software will be useful to the community as well as for similar programming practicals at other universities.

Bioinformatics

An Improved Genome Assembly of Azadirachta indica A. Juss.

Neem (Azadirachta indica A. Juss.), an evergreen tree of the Meliaceae family, is known for its medicinal, cosmetic, pesticidal and insecticidal properties. We had previously sequenced and published the draft genome of the plant, using mainly short read sequencing data. In this report, we present an improved genome assembly generated using additional short reads from Illumina and long reads from Pacific Biosciences SMRT sequencer. We assembled short reads and error-corrected long reads using Platanus, an assembler designed to perform well for heterozygous genomes. The updated genome assembly (v2.0) yielded 3- and 3.5-fold increase in N50 and N75, respectively; 2.6-fold decrease in the total number of scaffolds; 1.25-fold increase in the number of valid transcriptome alignments; 13.4-fold less mis-assembly and 1.85-fold increase in the percentage repeat, over the earlier assembly (v1.0). The current assembly also maps better to the genes known to be involved in the terpenoid biosynthesis pathway. Together, the data represents an improved assembly of the A. indica genome.\n\nThe raw data described in this manuscript are submitted to the NCBI Short Read Archive under the accession numbers SRX1074131, SRX1074132, SRX1074133, and SRX1074134 (SRP013453).

Bioinformatics

iCOBRA: open, reproducible, standardized and live method benchmarking

We present iCOBRA, a flexible general-purpose web-based application and accompanying R package to evaluate, compare and visualize the performance of methods for estimation or classification when ground truth is available. iCOBRA is interactive, can be run locally or remotely and generates customizable, publication-ready graphics. To facilitate open, reproducible and standardized method comparisons, expanding as new innovations are made, we encourage the community to provide benchmark results in a standard format.

Bioinformatics

nihexporter: an R package for NIH funding data

11.1 MotivationThe National Institutes of Health (NIH) is the major source of federal funding for biomedical research in the United States. Analysis of past and current NIH funding can illustrate funding trends and identify productive research topics, but these analyses are conducted ad hoc by the institutes themselves and only provide a small glimpse of the available data. The NIH provides free access to funding data via NIH EXPORTER, but no tools have been developed to enable analysis of this data.\n\n1.2 ResultsWe developed the nihexporter R package, which provides access to NIH EXPORTER data. We used the package to develop several analysis vignettes that show funding trends across NIH institutes over 15 years and highlight differences in how institutes change their funding profiles. Investigators and institutions can use the package to perform self-studies of their own NIH funding.\n\n1.3 AvailabilityThe nihexporter R package can be installed via github.\n\n1.4 ImplementationThe nihexporter package is implemented in the R Statistical Computing Environment.\n\n1.5 ContactJay Hesselberth jay.hesselberth@gmail.com, University of Colorado School of Medicine

Bioinformatics

TopHat-Recondition: A post-processor for TopHat unmapped reads

SummaryTopHat is a popular spliced junction mapper for RNA sequencing data, and writes files in the BAM format - the binary version of the Sequence Alignment/Map (SAM) format. BAM is the standard exchange format for aligned sequencing reads, thus correct format implementation is paramount for software interoperability and correct analysis. However, TopHat writes its unmapped reads in a way that is not compatible with other software that implements the SAM/BAM format. We have developed TopHat-Recondition, a post-processor for TopHat unmapped reads that restores read information in the proper format. TopHat-Recondition thus enables downstream software to process the plethora of BAM files written by TopHat.\n\nAvailability and implementationTopHat-Recondition is implemented in Python using the Pysam library and is freely available under a 2-clause BSD license on GitHub: https://github.com/cbrueffer/tophat-recondition.\n\nContact: christian.brueffer@med.lu.se, lao.saal@med.lu.se

Bioinformatics

Avoidance of stochastic RNA interactions can be harnessed to control protein expression levels in bacteria and archaea

A critical assumption of gene expression analysis is that mRNA abundances broadly correlate with protein abundance, but these two are often imperfectly correlated. Some of the discrepancy can be accounted for by two important mRNA features: codon usage and mRNA secondary structure. We present a new global factor, called mRNA:ncRNA avoidance, and provide evidence that avoidance increases translational efficiency. We also demonstrate a strong selection for avoidance of stochastic mRNA:ncRNA interactions across prokaryotes, and that these have a greater impact on protein abundance than mRNA structure or codon usage. By generating synonymously variant green fluorescent protein (GFP) mRNAs with different potential for mRNA:ncRNA interactions, we demonstrate that GFP levels correlate well with interaction avoidance. Therefore, taking stochastic mRNA:ncRNA interactions into account enables precise modulation of protein abundance.

Bioinformatics

Factorbook Motif Pipeline: A de novo motif discovery and filtering web server for ChIP-seq peaks

SummaryHigh-throughput sequencing technologies such as ChIP-seq have deepened our understanding in many biological processes. De novo motif search is one of the key downstream computational analysis following the ChIP-seq experiments and several algorithms have been proposed for this purpose. However, most web-based systems do not perform independent filtering or enrichment analyses to ensure the quality of the discovered motifs. Here, we developed a web server Factorbook Motif Pipeline based on an algorithm used in analyzing ENCODE consortium ChIP-seq datasets. It performs comprehensive analysis on the set of peaks detected from a ChIP-seq experiments: (i) de novo motif discovery; (ii) independent composition and bias analyses and (iii) matching to the annotated motifs. The statistical tests employed in our pipeline provide a reliable measure of confidence as to how significant are the motifs reported in the discovery step.\n\nAvailabilityFactorbook Motif Pipeline source code is accessible through the following URL. https://github.com/joshuabhk/factorbook-motif-pipeline

Bioinformatics

ME-plot: A QC package for bisulfite sequencing reads

Summary: Bisulfite sequencing is the gold standard method for analyzing methylomes. However, the effect of quality control in bisulfite sequencing has not been studied extensively. We developed a package ME-Plot to detect the errors in bisulfite sequencing mapping data and produce higher quality methylation calls by trimming low quality portions of the reads. Our simulation results on both randomly generated reads where methylation status is known and real world data indicate that ME-plot can detect errors in methylation mapping procedures and suggest post-processing steps to reduce the errors. ME-Plot requires SAM/BAM files for its analysis. Availability: The python package is available at the following URL. https://github.com/joshuabhk/methylsuite

Bioinformatics