Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Improving intermolecular contact prediction through protein-protein interaction prediction using coevolutionary analysis with expectation-maximization

Predicting residue-residue contacts between interacting proteins is an important problem in bioinformatics. The growing wealth of sequence data can be used to infer these contacts through correlated mutation analysis on multiple sequence alignments of interacting homologs of the proteins of interest. This requires correct identification of pairs of interacting proteins for many species, in order to avoid introducing noise (i.e. non-interacting sequences) in the analysis that will decrease predictive performance. We have designed Ouroboros, a novel algorithm to reduce such noise in intermolecular contact prediction. Our method iterates between weighting proteins according to how likely they are to interact based on the correlated mutations signal, and predicting correlated mutations based on the weighted sequence alignment. We show that this approach accurately discriminates between protein interaction versus noninteraction and simultaneously improves the prediction of intermolecular contact residues compared to a naive application of correlated mutation analysis. Furthermore, the method relaxes the assumption of one-to-one interaction of previous approaches, allowing for the study of many-to-many interactions. Source code and test data are available at www.bif.wur.nl/

bioinformatics

COSSMO: Predicting Competitive Alternative Splice Site Selection using Deep Learning

MotivationAlternative splice site selection is inherently competitive and the probability of a given splice site to be used also depends strongly on the strength of neighboring sites. Here we present a new model named Competitive Splice Site Model (COSSMO), which explicitly models these competitive effects and predict the PSI distribution over any number of putative splice sites. We model an alternative splicing event as the choice of a 3 acceptor site conditional on a fixed upstream 5 donor site, or the choice of a 5 donor site conditional on a fixed 3 acceptor site. We build four different architectures that use convolutional layers, communication layers, LSTMS, and residual networks, respectively, to learn relevant motifs from sequence alone. We also construct a new dataset from genome annotations and RNA-Seq read data that we use to train our model.\n\nResultsCOSSMO is able to predict the most frequently used splice site with an accuracy of 70% on unseen test data, and achieve an R2 of 60% in modeling the PSI distribution. We visualize the motifs that COSSMO learns from sequence and show that COSSMO recognizes the consensus splice site sequences as well as many known splicing factors with high specificity.\n\nAvailabilityOur dataset is available from http://cossmo.deepgenomics.com.\n\nContactfrey@deepgenomics.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

ClusterMine: a Knowledge-integrated Clustering Approach based on Expression Profiles of Gene Sets

MotivationClustering analysis is essential for understanding complex biological data. In widely used methods such as hierarchical clustering (HC) and consensus clustering (CC), expression profiles of all genes are often used to assess similarity between samples for clustering. These methods output sample clusters, but are not able to provide information about which gene sets (functions) contribute most to the clustering. So interpretability of their results is limited. We hypothesized that integrating prior knowledge of annotated biological processes would not only achieve satisfying clustering performance but also, more importantly, enable potential biological interpretation of clusters.\n\nResultsHere we report ClusterMine, a novel approach that identifies clusters by assessing functional similarity between samples through integrating known annotated gene sets, e.g., in Gene Ontology. In addition to outputting cluster membership of each sample as conventional approaches do, it outputs gene sets that are most likely to contribute to the clustering, a feature facilitating biological interpretation. Using three cancer datasets, two single cell RNA-sequencing based cell differentiation datasets, one cell cycle dataset and two datasets of cells of different tissue origins, we found that ClusterMine achieved similar or better clustering performance and that top-scored gene sets prioritized by ClusterMine are biologically relevant.\n\nImplementation and availabilityClusterMine is implemented as an R package and is freely available at: www.genemine.org/clustermine.php\n\nContactjxwang@csu.edu.cn\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics

PopPhy-CNN: A Phylogenetic Tree Embedded Architecture for Convolution Neural Networks for Metagenomic Data

MotivationAccurate prediction of the host phenotype from a metgenomic sample and identification of the associated bacterial markers are important in metagenomic studies. We introduce PopPhy-CNN, a novel convolutional neural networks (CNN) learning architecture that effectively exploits phylogentic structure in microbial taxa. PopPhy-CNN provides an input format of 2D matrix created by embedding the phylogenetic tree that is populated with the relative abundance of microbial taxa in a metagenomic sample. This conversion empowers CNNs to explore the spatial relationship of the taxonomic annotations on the tree and their quantitative characteristics in metagenomic data.\n\nResultsPopPhy-CNN is evaluated using three metagenomic datasets of moderate size. We show the superior performance of PopPhy-CNN compared to random forest, support vector machines, LASSO and a baseline 1D-CNN model constructed with relative abundance microbial feature vectors. In addition, we design a novel scheme of feature extraction from the learned CNN models and demonstrate the improved performance when the extracted features are used to train support vector machines.\n\nConclusionPopPhy-CNN is a novel deep learning framework for the prediction of host phenotype from metagenomic samples. PopPhy-CNN can efficiently train models and does not require excessive amount of data. PopPhy-CNN facilities not only retrieval of informative microbial taxa from the trained CNN models but also visualization of the taxa on the phynogenetic tree.\n\nContactyagndai@uic.edu\n\nAvailabilitySource code is publicly available at https://github.com/derekreiman/PopPhy-CNN\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

GITAR: An open source tool for analysis and visualization of Hi-C data

Interactions between chromatin segments play a large role in functional genomic assays and developments in genomic interaction detection methods have shown interacting topological domains within the genome. Among these methods, Hi-C plays a key role. Here, we present GITAR (Genome Interaction Tools and Resources), a software to perform a comprehensive Hi-C data analysis, including data preprocessing, normalization, visualization and topologically associated domains (TADs) analysis. GITAR is composed of two main modules: 1) HiCtool, a Python library to process and visualize Hi-C data, including TADs analysis and 2) Processed data library, a large collection of human and mouse datasets processed using HiCtool. HiCtool leads the user step-by-step through a pipeline which goes from the raw Hi-C data to the computation, visualization and optimized storage of intra-chromosomal contact matrices and topological domain coordinates. A large collection of standardized processed data allows to compare different datasets in a consistent way and it saves time of work to obtain data for visualization or additional analyses. GITAR enables users without any programming or bioinformatic expertise to work with Hi-C data and it is freely available for the public at http://genomegitar.org as an open source software.

bioinformatics

psichomics: graphical application for alternative splicing quantification and analysis

Alternative pre-mRNA splicing generates functionally distinct transcripts from the same gene and is involved in the control of multiple cellular processes, with its dysregulation being associated with a variety of pathologies. The advent of next-generation sequencing has enabled global studies of alternative splicing in different physiological and disease contexts. However, current bioinformatics tools for alternative splicing analysis from RNA-seq data are not user-friendly, disregard available exon-exon junction quantification or have limited downstream analysis features. To overcome such limitations, we have developed psichomics, an R package with an intuitive graphical interface for alternative splicing quantification and downstream dimensionality reduction, differential splicing and gene expression and survival analyses based on The Cancer Genome Atlas, the Genotype-Tissue Expression project and user-provided data. These integrative analyses can also incorporate clinical and molecular sample-associated features. We successfully used psichomics to reveal alternative splicing signatures specific to stage I breast cancer and associated novel putative prognostic factors.

bioinformatics

hts-nim: scripting high-performance genomic analyses

MotivationExtracting biological insight from genomic data inevitably requires custom software. In many cases, this is accomplished with scripting languages, owing to their accessibility and brevity. Unfortunately, the ease of scripting languages typically comes at a substantial performance cost that is especially acute with the scale of modern genomics datasets.\n\nResultsWe present hts-nim, a high-performance library written in the Nim programming language that provides a simple, scripting-like syntax without sacrificing performance.\n\nAvailabilityhts-nim is available at https://github.com/brentp/hts-nim and the example tools are at https://github.com/brentp/hts-nim-tools both under the MIT license.\n\nContactbpederse@gmail.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Transfer learning for biomedical named entity recognition with neural networks.

MotivationThe explosive increase of biomedical literature has made information extraction an increasingly important tool for biomedical research. A fundamental task is the recognition of biomedical named entities in text (BNER) such as genes/proteins, diseases, and species. Recently, a domain-independent method based on deep learning and statistical word embeddings, called long short-term memory network-conditional random field (LSTM-CRF), has been shown to outperform state-of-the-art entity-specific BNER tools. However, this method is dependent on gold-standard corpora (GSCs) consisting of hand-labeled entities, which tend to be small but highly reliable. An alternative to GSCs are silver-standard corpora (SSCs), which are generated by harmonizing the annotations made by several automatic annotation systems. SSCs typically contain more noise than GSCs but have the advantage of containing many more training examples. Ideally, these corpora could be combined to achieve the benefits of both, which is an opportunity for transfer learning. In this work, we analyze to what extent transfer learning improves upon state-of-the-art results for BNER.\n\nResultsWe demonstrate that transferring a deep neural network (DNN) trained on a large, noisy SSC to a smaller, but more reliable GSC significantly improves upon state-of-the-art results for BNER. Compared to a state-of-the-art baseline evaluated on 23 GSCs covering four different entity classes, transfer learning results in an average reduction in error of approximately 11%. We found transfer learning to be especially beneficial for target data sets with a small number of labels (approximately 6000 or less).\n\nAvailability and implementationSource code for the LSTM-CRF is available at https://github.com/Franck-Dernoncourt/NeuroNER/ and links to the corpora are available at https://github.com/BaderLab/Transfer-Learning-BNER-Bioinformatics-2018/.\n\nContactjohn.giorgi@utoronto.ca\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

TFmapper: A tool for searching putative factors regulating gene expression using ChIP-seq data

BackgroundNext-generation sequencing coupled to chromatin immunoprecipitation (ChIP-seq), DNase I hypersensitivity (DNase-seq) and the transposase-accessible chromatin assay (ATAC-seq) has generated enormous amounts of data, markedly improved our understanding of the transcriptional and epigenetic control of gene expression. To take advantage of the availability of such datasets and provide clues on what factors, including transcription factors, epigenetic regulators and histone modifications, potentially regulates the expression of a gene of interest, a tool for simultaneous queries of multiple datasets using symbols or genomic coordinates as search terms is needed.\n\nResultsIn this study, we annotated the peaks of thousands of ChIP-seq datasets generated by ENCODE project, or ChIP-seq/DNase-seq/ATAC-seq datasets deposited in Gene Expression Omnibus and curated by CistromeDB; We built a MySQL database called TFmapper containing the annotations and associated metadata, allowing users without bioinformatics expertise to search across thousands of datasets to identify factors targeting a genomic region/gene of interest in a specified sample through a web interface. Users can also visualize multiple peaks in genome browsers and download the corresponding sequences.\n\nConclusionTFmapper will help users explore the vast amount of publicly available ChIP-seq/DNase-seq/ATAC-seq data, and perform integrative analyses to understand the regulation of a gene of interest. The web server is freely accessible at http://www.tfmapper.org/.

bioinformatics

GeneQC: A quality control tool for gene expression estimation based on RNA-sequencing reads mapping

MotivationOne of the main benefits of using modern RNA-sequencing (RNA-Seq) technology is the more accurate gene expression estimations compared with previous generations of expression data, such as the microarray. However, numerous issues can result in the possibility that an RNA-Seq read can be mapped to multiple locations on the reference genome with the same alignment scores, which occurs in plant, animal, and metagenome samples. Such a read is so-called a multiple-mapping read (MMR). The impact of these MMRs is reflected in gene expression estimation and all downstream analyses, including differential gene expression, functional enrichment, etc. Current analysis pipelines lack the tools to effectively test the reliability of gene expression estimations, thus are incapable of ensuring the validity of all downstream analyses.\n\nResultsOur investigation into 95 RNA-Seq datasets from seven species (totaling 1,951GB) indicates an average of roughly 22% of all reads are MMRs for plant and animal species. Here we present a tool called GeneQC (Gene expression Quality Control), which can accurately estimate the reliability of each genes expression level. The underlying algorithm is designed based on extracted genomic and transcriptomic features, which are then combined using elastic-net regularization and mixture model fitting to provide a clearer picture of mapping uncertainty for each gene. GeneQC allows researchers to determine reliable expression estimations and conduct further analysis on the gene expression that is of sufficient quality. This tool also enables researchers to investigate continued re-alignment methods to determine more accurate gene expression estimates for those with low reliability.\n\nAvailabilityGeneQC is freely available at http://bmbl.sdstate.edu/GeneQC/home.html.\n\nContactqin.ma@sdstate.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

GLUE: A flexible software system for virus sequence data

Virus genome sequences, generated in ever-higher volumes, can provide new scientific insights and inform our responses to epidemics and outbreaks. To facilitate interpretation, such data must be organised and processed within scalable computing resources that encapsulate virology expertise. GLUE (Genes Linked by Underlying Evolution) is a data-centric bioinformatics environment for building such resources. Its flexible design emphasises applicability to different viruses and to diverse needs within research, clinical or public health contexts. A sequence data resource for hepatitis C virus (HCV) with clinical and research applications is presented as a case study.

bioinformatics

SMuRF: a novel tool to identify genomic regions enriched for somatic point mutations

MotivationSingle Nucleotide Variants (SNVs), including somatic point mutations and Single Nucleotide Polymorphisms (SNPs), in noncoding cis-regulatory elements (CREs) can affect gene regulation and lead to disease development (Zhou et al., 2016; Zhang et al., 2014). Others have previously developed methods to identify important clusters of somatic point mutations based on proximity (Weinhold et al., 2014) or the enrichment of inherited risk-SNPs at CREs (Ahmed et al., 2017). Here, we present SMuRF (Significantly Mutated Region Finder), a user-friendly command-line tool to identify these significantly mutated regions from user-defined genomic intervals and SNVs.\n\nResultsSMuRF identified 72 significantly mutated CREs in liver cancer, including known mutated gene promoters as well as previously unreported regions.\n\nAvailabilityThe source code for SMuRF is open-source and freely available on GitHub (https://github.com/LupienLabOrganization/SMuRF) under the GNU GPLv3 license. SMuRF is implemented in Bash and R; it runs on any platform with Bash ([≥]4.1.2), R ([≥]3.3.0) and BEDTools ([≥]2.26.0). It requires the following R packages: GenomicRanges, gtools, gplots, ggplot2, data.table, psych, and dplyr.\n\nSupplementary InformationSupplementary information available at Bioinformatics online.\n\nContactpaul.guilhamon@uhnresearch.ca; mlupien@uhnres.utoronto.ca

bioinformatics

Crossbrowse: A versatile genome browser for visualizing comparative experimental data

The recent beyond-exponential growth in diverse collections of deep sequencing datasets creates enormous opportunities for discovery, concomitant with new challenges for displaying and interpreting these data. Notably, the availability of scores of whole genome sequences in multiple species clades enables comparative studies of functional elements. However, current genome browsers do not permit effective visualization of multigenome experimental data. Here, we present CrossBrowse, a standalone desktop application for displaying and browsing cross-species genomic datasets. We utilize data standards and graphic representation of popular browsers, and incorporate an intuitive graphical visualization of genome synteny that facilitates and drives human interrogation of comparative data. Our platform permits users with minimal informatics capacity to select arbitrary sets of genomes for display, upload and configure multiple datasets, and interact with vertebrate-sized genomic datasets in real-time. We illustrate the utility of CrossBrowse with interrogation of comparative invertebrate and mammalian datasets that provide insights into diverse aspects of transcriptional and post-transcriptional regulation. Of note, we show examplars of both preservation and divergence of functional elements that cannot be inferred from sequence alignments alone. Moreover, we demonstrate how inspection of primary data using CrossBrowse exposes an artifact in a typical strategy for assigning species-specific functional elements, and drives the implementation of an improved computational strategy. We anticipate that CrossBrowse will greatly foster user-based discovery within multispecies genomic datasets, and inform their bioinformatic interpretation.

bioinformatics

SpiceRx: an integrated resource for the health impacts of culinary spices and herbs

Spices and herbs are key dietary ingredients used in cuisines across the world. They have been reported to be of medicinal value for a wide variety of diseases through a large body of biomedical investigations. Bioactive phytochemicals in these plant products form the basis of their therapeutic potential as well as adverse effects. A systematic compilation of empirical data involving these aspects of culinary spices and herbs could help unravel molecular mechanisms underlying their effects on health.\n\nSpiceRx provides a platform for exploring the health impact of spices and herbs used in food preparations through a structured database of tripartite relationships with their phytochemicals and disease associations. Starting with an extensive dictionary of culinary spices and herbs, their disease associations were text mined from MEDLINE, the largest database of biomedical abstracts, assisted with manual curation. This information was further combined with spice-phytochemical and phytochemical-disease associations. SpiceRx is an integrated repertoire of evidence-based knowledge pertaining to the health impacts of culinary spices and herbs, and facilitates their disease-specific culinary recommendations as well as exploration of molecular mechanisms underlying their health effects.\n\nAvailability and ImplementationSpiceRx is available at http://cosylab.iiitd.edu.in/spicerx and supports all modern browsers. SpiceRx is implemented with Python web development framework Django and relational database PostgreSQL; the front-end was built using HTML, CSS, JavaScript, AJAX, jQuery, JSME Molecular Editor, Bootstrap, Jmol, DataTables and Google Charts.\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

segment_liftover: a Python tool to convert segments between genome assemblies

The process of assembling a species reference genome may be performed in a number of iterations, with subsequent genome assemblies differing in the coordinates of mapped elements. The conversion of genome coordinates between different assemblies is required for many integrative and comparative studies. While currently a number of bioinformatics tools are available to accomplish this task, most of them are tailored towards the conversion of single genome coordinates. When converting the boundary positions of segments spanning larger genome regions, segments may be mapped into smaller subsegments if the original segments continuity is disrupted in the target assembly. Such a conversion may lead to a relevant degree of data loss in some circumstances such as copy number variation (CNV) analysis, where the quantitative representation of a genomic region takes precedence over base-specific accuracy. segment_liftover aims at continuity-preserving remapping of genome segments between assemblies and provides features such as approximate locus conversion, automated batch processing and comprehensive logging to facilitate processing of datasets containing large numbers of structural genome variation data.

bioinformatics

In Silico Processing of the Complete CRISPR-Cas Spacer Space for Identification of PAM Sequences

Despite extensive exploration of the diversity of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) systems, biological applications have been mostly confined to Class 2 systems, specifically the Cas9 and Cas12 (formerly Cpf1) single effector proteins. A key limitation of exploring and utilizing other CRISPR-Cas systems with unique functionalities, particularly Class I types and their multi-protein effector complex, is the knowledge of the systems protospacer adjacent motif (PAM) sequence identity. In this work, we developed a systematic pipeline, named CASPERpam, that enables us to comprehensively assess the PAM sequences of all the available CRISPR-Cas systems in the NCBI database of bacterial genomes. The CASPERpam analysis revealed that within the 30,389 assemblies previously screen for CRISPR arrays, there exists 26,364 spacers that match somewhere in the viral, bacterial, and plasmid databases of NCBI, using the constraints of 95% sequence identity and 95% sequence coverage for blast hits. When grouping these results by species, we were able to identify putative PAM sequences for 1,049 among 1,493 unique species. The remaining species either have insufficient data or an undetermined result from the analysis. Finally, we were able to infer certain design principles that are relevant for understanding PAM diversity and a baseline for further experimental studies including PAM assays. We envision CASPERpam is a useful bioinformatic tool for understanding and harnessing the diversity of CRISPR systems.

bioinformatics

Knomics-Biota - a system for exploratory analysis of human gut microbiota data

SummaryMetagenomic surveys of human microbiota are becoming increasingly widespread in academic research as well as in food and pharmaceutical industries and clinical context. Intuitive tools for exploration of experimental data are of high interest to researchers. Knomics-Biota is a Web-based resource for exploratory analysis of human gut metagenomes. Users can generate analytical reports that correspond to common experimental schemes (like case-control study or paired comparison). Statistical analysis and visualizations of microbiota composition are provided in association with the external factors and in the context of thousands of publicly available datasets.\n\nAvailability and ImplementationThe Web-service is available at https://biota.knomics.ru.\n\nContactanna.popenko@knomics.ru or a.tyakht@gmail.com.\n\nSupplementary informationSupplementary figures are available at Bioinformatics online.

bioinformatics

Protease target prediction via matrix factorization

MotivationProtein cleavage is an important cellular event, involved in a myriad of processes, from apoptosis to immune response. Bioinformatics provides in silico tools, such as machine learning-based models, to guide target discovery. State-of-the-art models have a scope limited to specific protease families (such as Caspases), and do not explicitly include biological or medical knowledge (such as the hierarchical protein domain similarity, or gene-gene interactions). To fill this gap, we present a novel approach for protease target prediction based on data integration.\n\nResultsBy representing protease-protein target information in the form of relational matrices, we design a model that: (a) is general, i.e., not limited to a single protease family; and (b) leverages on the available knowledge, managing extremely sparse data from heterogeneous data sources, including primary sequence, pathways, domains, and interactions from nine databases. When compared to other algorithms on test data, our approach provides a better performance even for models specifically focusing on a single protease family.\n\nAvailabilityhttps://gitlab.com/smarini/MaDDA/ (Matlab code and utilized data.)\n\nContactsmarini@med.umich.edu, or takutsu@kuicr.kyoto-u.ac.jp

bioinformatics