Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

A learning-based framework for miRNA-disease association prediction using neural networks

MotivationA microRNA (miRNA) is a type of non-coding RNA, which plays important roles in many bio-logical processes. Lots of studies have shown that miRNAs were implicated in human diseases, indicating that miRNAs might be potential biomarkers for various types of diseases. Therefore, it is important to reveal the relationships between miRNAs and diseases/phenotypes.\n\nResultsWe proposed a novel learning-based framework, MDA-CNN, for miRNA-disease association identification. The model first captures richer interaction features between diseases and miRNAs based on a three-layer network with an additional gene layer. An auto-encoder is then employed to identify the essential feature combination for each pair of miRNA and disease automatically. A final label is given by a convolutional neural network taking the reduced feature representation as input. The evaluation results showed that the proposed framework outperformed some state-of-the-art approaches in a large margin on both tasks of miRNA-disease associations prediction and miRNA-phenotype associations prediction.\n\nAvailabilityThe software will be available after the manuscript is accepted.\n\nContactjiajiepeng@nwpu.edu.cn;zywei@fudan.edu.cn;shang@nwpu.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

VarQ: a tool for the structural analysis of Human Protein Variants

Understanding the functional effect of Single Amino acid Substitutions (SAS), derived from the occurrence of single nucleotide variants (SNVs), and their relation to disease development is a major issue in clinical genomics. Even though there are several bioinformatic algorithms and servers that predict if a SAS can be pathogenic or not they give little or non-information on the actual effect on the protein function. Moreover, many of these algorithms are able to predict an effect that no necessarily translates directly into pathogenicity. VarQ Web Server is an online tool that given an UniProt id automatically analyzes known and user provided SAS for their effect on protein activity, folding, aggregation and protein interactions among others. VarQ assessment was performed over a set of previously manually curated variants, showing its ability to correctly predict the phenotypic outcome and its underlying cause. This resource is available online at http://varq.qb.fcen.uba.ar/.\n\nContact: lradusky@qb.fcen.uba.ar\n\nSupporting Information & Tutorials may be found in the webpage of the tool.

bioinformatics

VIGA: a sensitive, precise and automatic de novo VIral Genome Annotator.

Viral (meta)genomics is a rapidly growing field of study that is hampered by an inability to annotate the majority of viral sequences; therefore, the development of new bioinformatic approaches is very important. Here, we present a new automatic de novo genome annotation pipeline, called VIGA, to annotate prokaryotic and eukaryotic viral sequences from (meta)genomic studies. VIGA was benchmarked on a database of known viral genomes and a viral metagenomics case study. VIGA generated the most accurate outputs according to the number of coding sequences and their coordinates, outputs also had a lower number of non-informative annotations compared to other programs.

bioinformatics

Harnessing Empirical Bayes and Mendelian Segregation for Genotyping Autopolyploids from Messy Sequencing Data

Detecting and quantifying the differences in individual genomes (i.e. genotyping), plays a fundamental role in most modern bioinformatics pipelines. Many scientists now use reduced representation next-generation sequencing (NGS) approaches for genotyping. Genotyping diploid individuals using NGS is a well-studied field and similar methods for polyploid individuals are just emerging. However, there are many aspects of NGS data, particularly in polyploids, that remain unexplored by most methods. We provide two main contributions in this paper: (i) We draw attention to, and then model, common aspects of NGS data: sequencing error, allelic bias, overdispersion, and outlying observations. (ii) Many datasets feature related individuals, and so we use the structure of Mendelian segregation to build an empirical Bayes approach for genotyping polyploid individuals. We assess the accuracy of our method in simulations and apply it to a dataset of hexaploid sweet potatoes (Ipomoea batatas). An R package implementing our method is available at https://github.com/dcgerard/updog.

bioinformatics

Tychus: a whole genome sequencing pipeline for assembly, annotation and phylogenetics of bacterial genomes

SummaryTychus is a tool that allows researchers to perform massively parallel whole genome sequence (WGS) analysis with the goal of producing a high confidence and comprehensive description of the bacterial genome. Key features of the Tychus pipeline include the assembly, annotation, alignment, variant discovery and phylogenetic inference of large numbers of WGS isolates in parallel using open-source bioinformatics tools and virtualization technology. All prerequisite tools and dependencies come packaged together in a single suite that can be easily downloaded and installed on Linux and Mac operating systems.\n\nAvailabilityTychus is freely available as an open-source package under the MIT license, and can be downloaded via GitHub (https://github.com/Abdo-Lab/Tychus).\n\nContactzaid.abdo@colostate.edu

bioinformatics

IRIS-DGE: An integrated RNA-seq data analysis and interpretation system for differential gene expression

MotivationNext-Generation Sequencing has made available much more large-scale genomic and transcriptomic data. Studies with RNA-sequencing (RNA-seq) data typically involve generation of gene expression profiles that can be further analyzed, many times involving differential gene expression (DGE). This process enables comparison across samples of two or more factor levels. A recurring issue with DGE analyses is the complicated nature of the comparisons to be made, in which a variety of factor combinations, pairwise comparisons, and main or blocked main effects need to be tested.\n\nResultsHere we present a tool called IRIS-DGE, which is a server-based DGE analysis tool developed using Shiny. It provides a straightforward, user-friendly platform for performing comprehensive DGE analysis, and crucial analyses that help design hypotheses and to determine key genomic features. IRIS-DGE integrates the three most commonly used R-based DGE tools to determine differentially expressed genes (DEGs) and includes numerous methods for performing preliminary analysis on user-provided gene expression information. Additionally, this tool integrates a variety of visualizations, in a highly interactive manner, for improved interpretation of preliminary and DGE analyses.\n\nAvailabilityIRIS-DGE is freely available at http://bmbl.sdstate.edu/IRIS/.\n\nContactqin.ma@sdstate.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

MechRNA: prediction of lncRNA mechanisms from RNA-RNA and RNA-protein interactions

MotivationLong non-coding RNAs (lncRNAs) are defined as transcripts longer than 200 nucleotides that do not get translated into proteins. Often these transcripts are processed (spliced, capped, polyadenylated) and some are known to have important biological functions. However, most lncRNAs have unknown or poorly understood functions. Nevertheless, because of their potential role in cancer, lncRNAs are receiving a lot of attention, and the need for computational tools to predict their possible mechanisms of action is more than ever. Fundamentally, most of the known lncRNA mechanisms involve RNA-RNA and/or RNA-protein interactions. Through accurate predictions of each kind of interaction and integration of these predictions, it is possible to elucidate potential mechanisms for a given lncRNA.\n\nApproachHere we introduce MechRNA, a pipeline for corroborating RNA-RNA interaction prediction and protein binding prediction for identifying possible lncRNA mechanisms involving specific targets or on a transcriptome-wide scale. The first stage uses a version of IntaRNA2 with added functionality for efficient prediction of RNA-RNA interactions with very long input sequences, allowing for large-scale analysis of lncRNA interactions with little or no loss of optimality. The second stage integrates protein binding information pre-computed by GraphProt, for both the lncRNA and the target. The final stage involves inferring the most likely mechanism for each lncRNA/target pair. This is achieved by generating candidate mechanisms from the predicted interactions, the relative locations of these interactions and correlation data, followed by selection of the most likely mechanistic explanation using a combined p-value.\n\nResultsWe applied MechRNA on a number of recently identified cancer-related lncRNAs (PCAT1, PCAT29, ARLnc1) and also on two well-studied lncRNAs (PCA3 and 7SL). This led to the identification of hundreds of high confidence potential targets for each lncRNA and corresponding mechanisms. These predictions include the known competitive mechanism of 7SL with HuR for binding on the tumor suppressor TP53, as well as mechanisms expanding what is known about PCAT1 and ARLn1 and their targets BRCA2 and AR, respectively. For PCAT1-BRCA2, the mechanism involves competitive binding with HuR, which we confirmed using HuR immunoprecipitation assays.\n\nAvailabilityMechRNA is available for download at https://bitbucket.org/compbio/mechrna\n\nContactbackofen@informatik.uni-freiburg.de, cenksahi@indiana.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Gene2Vec: Distributed Representation of Genes Based on Co-Expression

Existing functional description of genes are categorical, discrete, and mostly through manual process. In this work, we explore the idea of gene embedding, distributed representation of genes, in the spirit of word embedding. From a pure data-driven fashion, we trained a 300 dimension vector representation of all human genes, using gene co-expression patterns in 984 data sets from the GEO databases. These vectors capture functional relatedness of genes in terms of recovering known pathways - the average inner product (similarity) of genes within a pathway is 1.68X greater than that of random genes. Using t-SNE, we produced a gene co-expression map that shows local concentrations of tissue specific genes. We also illustrated the usefulness of the embedded gene vectors, laden with rich information on gene co-expression patterns, in tasks such as gene-gene interaction prediction. Overall, we believe that this distributed representation of genes may be useful for more bioinformatics applications.

bioinformatics

snpAD: An ancient DNA genotype caller

MotivationThe study of ancient genomes can elucidate the evolutionary past. However, analyses are complicated by base-modifications in ancient DNA molecules that result in errors in DNA sequences. These errors are particularly common near the ends of sequences and pose a challenge for genotype calling.\n\nResultsI describe an expectation-maximization algorithm that estimates genotype frequencies and errors along sequences to allow for accurate genotype calling from ancient sequences. The implementation of this method, called snpAD, performs well on high-coverage ancient data, as shown by simulations and by subsampling the data of a high-coverage Neandertal genome. Although estimates for low-coverage genomes are less accurate, I am able to derive approximate estimates of heterozygosity from several low-coverage Neandertals. These estimates show that low heterozygosity, compared to modern humans, was common among Neandertals.\n\nAvailabilityThe C++ code of snpAD is freely available at http://bioinf.eva.mpg.de/snpAD/\n\nContactpruefer@eva.mpg.de\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Porter 5: state-of-the-art ab initio prediction of protein secondary structure in 3 and 8 classes

MotivationAlthough Secondary Structure Predictors have been developed for more than 60 years, current ab initio methods have still some way to go to reach their theoretical limits. Moreover, the continuous effort towards harnessing ever increasing data sets and more sophisticated, deeper Machine Learning techniques, has not come to an end.\n\nResultsHere we present Porter 5, the last release of one of the best performing ab initio secondary structure predictor. Version 5 achieves 84% accuracy (84% SOV) when tested on 3 classes, and 73% accuracy (82% SOV) on 8 classes, on a large independent set, significantly outperforming all the most recent ab initio predictors we have tested.\n\nAvailabilityThe web and standalone versions of Porter5 are available at http://distilldeep.ucd.ie/porter.\n\nContactname@bio.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

TBtools, a Toolkit for Biologists integrating various HTS-data handling tools with a user-friendly interface

The rapid development of high-throughput sequencing (HTS) techniques has led biology into the big-data era. Data analyses using various bioinformatics tools rely on programming and command-line environments, which are challenging and time-consuming for most wet-lab biologists. Here, we present TBtools (a Toolkit for Biologists integrating various biological data handling tools), a stand-alone software with a user-friendly interface. The toolkit incorporates over 100 functions, which are designed to meet the increasing demand for big-data analyses, ranging from bulk sequence processing to interactive data visualization. A wide variety of graphs can be prepared in TBtools, with a new plotting engine ("JIGplot") developed to maximum their interactive ability, which allows quick point-and-click modification to almost every graphic feature. TBtools is a platform-independent software that can be run under all operating systems with Java Runtime Environment 1.6 or newer. It is freely available to non-commercial users at https://github.com/CJ-Chen/TBtools/releases.

bioinformatics

Understanding the limit of open search in the identification of peptides with post-translational modifications -- A simulation-based study

MotivationAnalyzing tandem mass spectrometry data to recognize peptides in a sample is the fundamental task in computational proteomics. Traditional peptide identification algorithms perform well when identifying unmodified peptides. However, when peptides have post-translational modifications (PTMs), these methods cannot provide satisfactory results. Recently, Chick et al., 2015 and Yu et al., 2016 proposed the spectrum-based and tag-based open search methods, respectively, to identify peptides with PTMs. While the performance of these two methods is promising, the identification results vary greatly with respect to the quality of tandem mass spectra and the number of PTMs in peptides. This motivates us to systematically study the relationship between the performance of open search methods and quality parameters of tandem mass spectrum data, as well as the number of PTMs in peptides.\n\nResultsThrough large-scale simulations, we obtain the performance trend when simulated tandem mass spectra are of different quality. We propose an analytical model to describe the relationship between the probability of obtaining correct identifications and the spectrum quality as well as the number of PTMs. Based on the analytical model, we can quantitatively describe the necessary condition to effectively apply open search methods.\n\nAvailabilitySource codes of the simulation are available at http://bioinformatics.ust.hk/PST.html.\n\nContactboningli@ust.hk or eeyu@ust.hk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Integrative analysis of pharmacogenomics in major cancer cell line databases using CellMinerCDB

As precision medicine demands molecular determinants of drug response, CellMinerCDB provides (https://discover.nci.nih.gov/cellminercdb/) a web-based portal for multiple forms of pharmacological, molecular, and genomic analyses, unifying the richest cancer cell line datasets (NCI-60, NCI-SCLC, Sanger/MGH GDSC, and Broad CCLE/CTRP). CellMinerCDB enables genomic and pharmacological data queries for identifying pharmacogenomic determinants, drug signatures, and gene regulatory networks for researchers without requiring specialized bioinformatics support. It leverages overlaps of cell lines and tested drugs to allow assessment of data reproducibility. It builds on the complementarity and strength of each dataset. A panel of 41 drugs evaluated in parallel in the NCI-60 and GDSC is reported, supporting drug reproducibility across databases, repositioning of bisacodyl and acetalax for triple negative breast cancer, and identifying novel drug response determinants and genomic signatures for topoisomerase inhibitors and schweinfurthins in development. CellMinerCDB also allowed the identification of LIX1L as a novel mesenchymal gene regulating cellular migration and invasiveness.

bioinformatics

RobusTAD: A Tool for Robust Annotation ofTopologically Associating Domain Boundaries

MotivationTopologically Associating Domains (TADs) are chromatin structures that can be identified by analysis of Hi-C data. Tools currently available for TAD identification are sensitive to experimental conditions such as coverage, resolution and noise level.\n\nResultsHere, we present RobusTAD, a tool to score TAD boundaries in a manner that is robust to these parameters. In doing so, RobusTAD eases comparative analysis of TAD structures across multiple heterogeneous samples.\n\nAvailabilityRobusTAD is implemented in R and released under a GPL license. RobusTAD can be downloaded from https://github.com/rdali/RobusTAD and runs on any standard desktop computer.\n\nContactrola.dali@mail.mcgill.ca, blanchem@cs.mcgill.ca\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

A unifying framework for joint trait analysis under a non-infinitesimal model

MotivationA large proportion of risk regions identified by genome-wide association studies (GWAS) are shared across multiple diseases and traits. Understanding whether this clustering is due to sharing of causal variants or chance colocalization can provide insights into shared etiology of complex traits and diseases.\n\nResultsIn this work, we propose a flexible, unifying framework to quantify the overlap between a pair of traits called UNITY (Unifying Non-Infinitesimal Trait analYsis). We formulate a Bayesian generative model that relates the overlap between pairs of traits to GWAS summary statistic data under a non-infinitesimal genetic architecture underlying each trait. We propose a Metropolis-Hastings sampler to compute the posterior density of the genetic overlap parameters in this model. We validate our method through comprehensive simulations and analyze summary statistics from height and BMI GWAS to show that it produces estimates consistent with the known genetic makeup of both traits.\n\nAvailabilityThe UNITY software is made freely available to the research community at: https://github.com/bogdanlab/UNITY\n\nContactruthjohnson@ucla.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

iTOP: Inferring the Topology of Omics Data

MotivationIn biology, we are often faced with multiple datasets recorded on the same set of objects, such as multi-omics and phenotypic data of the same tumors. These datasets are typically not independent from each other. For example, methylation may influence gene expression, which may, in turn, influence drug response. Such relationships can strongly affect analyses performed on the data, as we have previously shown for the identification of biomarkers of drug response. Therefore, it is important to be able to chart the relationships between datasets.\n\nResultsWe present iTOP, a methodology to infera topology of relationships between datasets. We base this methodology on the RV coefficient, a measure of matrix correlation, which can be used to determine how much information is shared between two datasets. We extended the RV coefficient for partial matrix correlations, which allows the use of graph reconstruction algorithms, such as the PC algorithm, to infer the topologies. In addition, since multi-omics data often contain binary data (e.g. mutations), we also extended the RV coefficient for binary data. Applying iTOP to pharmacogenomics data, we found that gene expression acts as a mediator between most other datasets and drug response: only proteomics clearly shares information with drug response that is not present in gene expression. Based on this result, we used TANDEM, a method for drug response prediction, to identify which variables predictive of drug response were distinct to either gene expression or proteomics.\n\nAvailabilityAn implementation of our methodology is available in the R package iTOP on CRAN. Additionally, an R Markdown document with code to reproduce all figures is provided as Supplementary Material.\n\nContacta.k.smilde@uva.nl and l.wessels@nki.nl\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

xAtlas: Scalable small variant calling across heterogeneous next-generation sequencing experiments

MotivationThe rapid development of next-generation sequencing (NGS) technologies has lowered the barriers to genomic data generation, resulting in millions of samples sequenced across diverse experimental designs. The growing volume and heterogeneity of these sequencing data complicate the further optimization of methods for identifying DNA variation, especially considering that curated highconfidence variant call sets commonly used to evaluate these methods are generally developed by reference to results from the analysis of comparatively small and homogeneous sample sets.\n\nResultsWe have developed xAtlas, an application for the identification of single nucleotide variants (SNV) and small insertions and deletions (indels) in NGS data. xAtlas is easily scalable and enables execution and retraining with rapid development cycles. Generation of variant calls in VCF or gVCF format from BAM or CRAM alignments is accomplished in less than one CPU-hour per 30x short-read human whole-genome. The retraining capabilities of xAtlas allow its core variant evaluation models to be optimized on new sample data and user-defined truth sets. Obtaining SNV and indels calls from xAtlas can be achieved more than 40 times faster than established methods while retaining the same accuracy.\n\nAvailabilityFreely available under a BSD 3-clause license at https://github.com/jfarek/xatlas.\n\nContactfarek@bcm.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Pan-cancer analysis of transcriptional metabolic dysregulation using The Cancer Genome Atlas

Understanding the levels of metabolic dysregulation in different disease settings is vital for the safe and effective incorporation of metabolism-targeted therapeutics in the clinic. Using transcriptomic data from 10,704 tumor and normal samples from The Cancer Genome Atlas, across 26 disease sites, we developed a novel bioinformatics pipeline that distinguishes tumor from normal tissues, based on differential gene expression for 114 metabolic pathways. This pathway dysregulation was confirmed in separate patient populations, further demonstrating the robustness of this approach. A bootstrapping simulation was then applied to assess whether these alterations were biologically meaningful, rather than expected by chance. We provide distinct examples of the types of analysis that can be accomplished with this tool to understand cancer specific metabolic dysregulation, highlighting novel pathways of interest in both common and rare disease sites. Utilizing a pathway mapping approach to understand patterns of metabolic flux, differential drug sensitivity, can accurately be predicted. Further, the identification of Master Metabolic Transcriptional Regulators, whose expression was highly correlated with pathway gene expression, explains why metabolic differences exist in different disease sites. We demonstrate these also have the ability to segregate patient populations and predict responders to different metabolism-targeted therapeutics.

bioinformatics