Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

BALLI: Bartlett-Adjusted Likelihood-based LInear Model Approach for Identifying Differentially Expressed Gene with RNA-seq Data

MotivationTranscriptomic profiles can improve our understanding of the phenotypic molecular basis of biological research, and many statistical methods have been proposed to identify differentially expressed genes under two or more conditions with RNA-seq data. However, statistical analyses with RNA-seq data often suffer from small sample sizes, and global variance estimates of RNA expression levels have been utilized as prior distributions for gene-specific variance estimates, making it difficult to generalize the methods to more complicated settings. We herein proposed a Bartlett-Adjusted Likelihood based LInear mixed model approach (BALLI) to analyze more complicated RNA-seq data. The proposed method estimates the technical and biological variances with a linear mixed effect model, with and without adjusting small sample bias using Bartletts corrections.\n\nResultsWe conducted extensive simulations to compare the performance of BALLI with those of existing approaches (edgeR, DESeq2, and voom). Results from the simulation studies showed that BALLI correctly controlled the type-1 error rates at the various nominal significance levels, and produced better statistical power and precision estimates than those of other competing methods in various scenarios. Furthermore, BALLI was robust to variation of library size. It was also successfully applied to Holstein milk yield data, illustrating its practical value.\n\nAvailability and ImplementationBALLI is implemented as R package and freely available at http://healthstat.snu.ac.kr/software/balli/.\n\nContactwon1@snu.ac.kr\n\nSupplementary InformationSupplementary data are available at Bioinformatics online

bioinformatics

Probabilistic variable-length segmentation of protein sequences for discriminative motif mining (DiMotif)and sequence embedding (ProtVecX)

In this paper, we present peptide-pair encoding (PPE), a general-purpose probabilistic segmentation of protein sequences into commonly occurring variable-length sub-sequences. The idea of PPE segmentation is inspired by the byte-pair encoding (BPE) text compression algorithm, which has recently gained popularity in subword neural machine translation. We modify this algorithm by adding a sampling framework allowing for multiple ways of segmenting a sequence. PPE segmentation steps can be learned over a large set of protein sequences (Swiss-Prot) or even a domain-specific dataset and then applied to a set of unseen sequences. This representation can be widely used as the input to any downstream machine learning tasks in protein bioinformatics. In particular, here, we introduce this representation through protein motif discovery and protein sequence embedding. (i) DiMotif: we present DiMotif as an alignment-free discriminative motif discovery method and evaluate the method for finding protein motifs in three different settings: (1) comparison of DiMotif with two existing approaches on 20 distinct motif discovery problems which are experimentally verified, (2) classification-based approach for the motifs extracted for integrins, integrin-binding proteins, and biofilm formation, and (3) in sequence pattern searching for nuclear localization signal. The DiMotif, in general, obtained high recall scores, while having a comparable F1 score with other methods in the discovery of experimentally verified motifs. Having high recall suggests that the DiMotif can be used for short-list creation for further experimental investigations on motifs. In the classification-based evaluation, the extracted motifs could reliably detect the integrins, integrin-binding, and biofilm formation-related proteins on a reserved set of sequences with high F1 scores. (ii) ProtVecX: we extend k-mer based protein vector (ProtVec) embedding to variable-length protein embedding using PPE sub-sequences. We show that the new method of embedding can marginally outperform ProtVec in enzyme prediction as well as toxin prediction tasks. In addition, we conclude that the embeddings are beneficial in protein classification tasks when they are combined with raw k-mer features.\n\nAvailabilityImplementations of our method will be available under the Apache 2 licence at http://llp.berkeley.edu/dimotif and http://llp.berkeley.edu/protvecx.

bioinformatics

SqueezeM, a fully automatic metagenomic analysis pipeline from reads to bins

The improvement of sequencing technologies has allowed the generalization of metagenomic sequencing, which has become a standard procedure for analysing the structure and functionality of microbiomes. The bioinformatic analysis of the sequencing results poses a challenge because it involves many different complex steps. SqueezeMeta is a full automatic pipeline for metagenomics/metatranscriptomics, covering all steps of the analysis. SqueezeMeta includes multi-metagenome support allowing the co-assembly of related metagenomes and the retrieval of individual genomes via binning procedures. SqueezeMeta features several unique characteristics: Co-assembly procedure or co-assembly of unlimited number of metagenomes via merging of individual assembled metagenomes, both with read mapping for estimation of the abundances of genes in each metagenome. It also includes binning and bin checking for retrieving individual genomes. Internal checks for the assembly and binning steps inform about the consistency of contigs and bins. Also, the results are stored in a mySQL database, where they can be easily exported and shared, and can be inspected anywhere using a flexible web interface allowing the easy creation of complex queries.\n\nWe illustrate the potential of SqueezeMeta by analyzing 32 gut metagenomes in a fully automatic way, allowing to retrieve several millions of genes and several hundreds of genomic bins.\n\nOne of the motivations in the development of SqueezeMeta was producing a software capable to run in small desktop computers, thus being amenable to all users and all settings. We were also able to co-assemble two of these metagenomes and complete the full analysis in less than one day using a simple laptop computer, illustrating the capacity of SqueezeMeta to run without high-performance computing infrastructure. SqueezeMeta is a complete system covering all steps in the analysis of metagenomes and metatranscriptomes, capable to work even in scarcity of computational resources. It is therefore adequate for in-situ, real time analysis of metagenomes produced by nanopore sequencing.\n\nSqueezeMeta can be downloaded from https://github.com/jtamames/SqueezeMeta

bioinformatics

Analyses of cancer data in the Genomic Data Commons Data Portal with new functionalities in the TCGAbiolinks R/Bioconductor package

The advent of Next Generation Sequencing (NGS) technologies has opened new perspectives in deciphering the genetic mechanisms underlying complex diseases. Nowadays, the amount of genomic data is massive and substantial efforts and new tools are required to unveil the information hidden in the data.\n\nThe Genomic Data Commons (GDC) Data Portal is a large data collection platform that includes different genomic studies included the ones from The Cancer Genome Atlas (TCGA) and the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) initiatives, accounting for more than 40 tumor types originating from nearly 30000 patients. Such platforms, although very attractive, must make sure the stored data are easily accessible and adequately harmonized. Moreover, they have the primary focus on the data storage in a unique place, and they do not provide a comprehensive toolkit for analyses and interpretation of the data. To fulfill this urgent need, comprehensive but easily accessible computational methods for integrative analyses of genomic data without renouncing a robust statistical and theoretical framework are needed. In this context, the R/Bioconductor package TCGAbiolinks was developed, offering a variety of bioinformatics functionalities. Here we introduce new features and enhancements of TCGAbiolinks in terms of i) more accurate and flexible pipelines for differential expression analyses, ii) different methods for tumor purity estimation and filtering, iii) integration of normal samples from the Genotype-Tissue-Expression (GTEx) platform iv) support for other genomics datasets, here exemplified by the TARGET data.\n\nEvidence has shown that accounting for tumor purity is essential in the study of tumorigenesis, as these factors promote confounding behavior regarding differential expression analysis. Henceforth, we implemented these filtering procedures in TCGAbiolinks. Moreover, a limitation of some of the TCGA datasets is the unavailability or paucity of corresponding normal samples. We thus integrated into TCGAbiolinks the possibility to use normal samples from the Genotype-Tissue Expression (GTEx) project, which is another large-scale repository cataloging gene expression from healthy individuals. The new functionalities are available in the TCGABiolinks v 2.8 and higher released in Bioconductor version 3.7.

bioinformatics

BioJupies: Automated Generation of Interactive Notebooks for RNA-seq Data Analysis in the Cloud

Interactive notebooks can make bioinformatics data analyses more transparent, accessible and reusable. However, creating notebooks requires computer programming expertise. Here we introduce BioJupies, a web server that enables automated creation, storage, and deployment of Jupyter Notebooks containing RNA-seq data analyses. Through an intuitive interface, novice users can rapidly generate tailored reports to analyze and visualize their own raw sequencing files, their gene expression tables, or fetch data from >5,500 published studies containing >250,000 preprocessed RNA-seq samples. Generated notebooks have executable code of the entire pipeline, rich narrative text, interactive data visualizations, and differential expression and enrichment analyses. The notebooks are permanently stored in the cloud and made available online through a persistent URL. The notebooks are downloadable, customizable, and can run within a Docker container. By providing an intuitive user interface for notebook generation for RNA-seq data analysis, starting from the raw reads, all the way to a complete interactive and reproducible report, BioJupies is a useful resource for experimental and computational biologists. BioJupies is freely available as a web-based application from: http://biojupies.cloud and as a Chrome extension from the Chrome Web Store.

bioinformatics

A deep learning approach to pattern recognition for short DNA sequences

MotivationInferring properties of biological sequences--such as determining the species-of-origin of a DNA sequence or the function of an amino-acid sequence--is a core task in many bioinformatics applications. These tasks are often solved using string-matching to map query sequences to labeled database sequences or via Hidden Markov Model-like pattern matching. In the current work we describe and assess an deep learning approach which trains a deep neural network (DNN) to predict database-derived labels directly from query sequences. ResultsWe demonstrate this DNN performs at state-of-the-art or above levels on a difficult, practically important problem: predicting species-of-origin from short reads of 16S ribosomal DNA. When trained on 16S sequences of over 13,000 distinct species, our DNN achieves read-level species classification accuracy within 2.0% of perfect memorization of training data, and produces more accurate genus-level assignments for reads from held-out species than k-mer, alignment, and taxonomic binning baselines. Moreover, our models exhibit greater robustness than these existing approaches to increasing noise in the query sequences. Finally, we show that these DNNs perform well on experimental 16S mock community dataset. Overall, our results constitute a first step towards our long-term goal of developing a general-purpose deep learning approach to predicting meaningful labels from short biological sequences. AvailabilityTensorFlow training code is available through GitHub (https://github.com/tensorflow/models/tree/master/research). Data in TensorFlow TFRecord format is available on Google Cloud Storage (gs://brain-genomics-public/research/seq2species/). Contactseq2species-interest@google.com Supplementary informationSupplementary data are available in a separate document.

bioinformatics

DeepImpute: an accurate, fast and scalable deep neural network method to impute single-cell RNA-Seq data

BackgroundSingle-cell RNA sequencing (scRNA-seq) offers new opportunities to study gene expression of tens of thousands of single cells simultaneously. However, a significant problem of current scRNA-seq data is the large fractions of missing values or \"dropouts\" in gene counts. Incorrect handling of dropouts may affect downstream bioinformatics analysis. As the number of scRNA-seq datasets grows drastically, it is crucial to have accurate and efficient imputation methods to handle these dropouts.\n\nMethodsWe present DeepImpute, a deep neural network based imputation algorithm. The architecture of DeepImpute efficiently uses dropout layers and loss functions to learn patterns in the data, allowing for accurate imputation.\n\nResultsOverall DeepImpute yields better accuracy than other publicly available scRNA-Seq imputation methods on experimental data, as measured by mean squared error or Pearsons correlation coefficient. Moreover, its efficient implementation provides significantly higher performance over the other methods as dataset size increases. Additionally, as a machine learning method, DeepImpute allows to use a subset of data to train the model and save even more computing time, without much sacrifice on the prediction accuracy.\n\nConclusionsDeepImpute is an accurate, fast and scalable imputation tool that is suited to handle the ever increasing volume of scRNA-seq data. The package is freely available at https://github.com/lanagarmire/DeepImpute

bioinformatics

Simulation-based approaches to characterize the effect of sequencing depth on the quantity and quality of metagenome-assembled genomes

We applied simulation-based approaches to characterize how microbial community structure influences the amount of sequencing effort to reconstruct metagenomes that are assembled from short read sequences. An initial analysis evaluated the quantity, completion, and contamination of complete-metagenome-assembled genome (complete-MAG) equivalents, a bioinformatic-pipeline normalized metric for MAG quantity, as a function of sequencing effort, on four preexisting sequence read datasets taken from a maize soil, an estuarine sediment, the surface ocean, and the human gut. These datasets were subsampled to varying degrees of completeness in order to simulate the effect of sequencing effort on MAG retrieval. Modeling suggested that sequencing efforts beyond what is typical in published experiments (1 to 10 Gbp) would generate diminishing returns in terms of MAG binning. A second analysis explored the theoretical relationship between sequencing effort and the proportion of available metagenomic DNA sequenced during a sequencing experiment as a function of community richness, evenness, and genome size. Simulations from this analysis demonstrated that while community richness and evenness influenced the amount of sequencing required to sequence a community metagenome to exhaustion, the effort necessary to sequence an individual genome to a target fraction of exhaustion was only dependent on the relative abundance of the corresponding organism and its genome size. A software tool, GRASE, was created to assist investigators further explore this relationship. Re-evaluation of the relationship between sequencing effort and binning success in the context of the relative abundance of genomes, as opposed to base pairs, provides a framework to design sequencing experiments based on the relative abundance of microbes in an environment rather than arbitrary levels of sequencing effort.

bioinformatics

In situ transcriptome characteristics are lost following culture adaptation of adult cardiac stem cells

Regenerative therapeutic approaches for myocardial diseases often involve adoptive transfer of stem cells expanded ex vivo. Prior studies indicate that cell culture conditions affect functional and phenotypic characteristics, but relationship(s) of cultured cells derived from freshly isolated populations and the heterogeneity of the cultured population remain poorly defined. Functional and phenotypic characteristics of adoptively donated cells will determine outcomes of interventional treatment for disease, necessitating characterization of the impact that ex vivo expansion has upon isolated stem cell populations. Single-cell RNA-Seq profiling (scRNA-Seq) was performed to determine consequences of culture expansion upon adult cardiac progenitor cells (CPCs) as well as relationships with other cell populations. Bioinformatic analyses reveal loss of identity marker genes in cultured CPCs while simultaneously acquiring thousands of additional genes. Cultured CPCs exhibited decreased transcriptome variability within their population relative to their freshly isolated cells. Findings were validated by comparative analyses using scRNA-Seq datasets of various cell types generated by multiple scRNA-Seq technology. Increased transcriptome diversity and decreased population heterogeneity in the cultured cell population relative to freshly isolated cells may help account for reported outcomes associated with experimental and clinical use of CPCs for treatment of myocardial injury.

bioinformatics

An open-source k-mer based machine learning tool for fast and accurate subtyping of HIV-1 genomes

For many disease-causing virus species, global diversity is clustered into a taxonomy of subtypes with clinical significance. In particular, the classification of infections among the subtypes of human immunodeficiency virus type 1 (HIV-1) is a routine component of clinical management, and there are now many classification algorithms available for this purpose. Although several of these algorithms are similar in accuracy and speed, the majority are proprietary and require laboratories to transmit HIV-1 sequence data over the network to remote servers. This potentially exposes sensitive patient data to unauthorized access, and makes it impossible to determine how classifications are made and to maintain the data provenance of clinical bioinformatic workflows. We propose an open-source supervised and alignment-free subtyping method (KO_SCPCAPAMERISC_SCPCAP) that operates on k-mer frequencies in HIV-1 sequences. We performed a detailed study of the accuracy and performance of subtype classification in comparison to four state-of-the-art programs. Based on our testing data set of manually curated real-world HIV-1 sequences (n = 2, 784), Kameris obtained an overall accuracy of 97%, which matches or exceeds all other tested software, with a processing rate of over 1,500 sequences per second. Furthermore, our fully standalone general-purpose software provides key advantages in terms of data security and privacy, transparency and reproducibility. Finally, we show that our method is readily adaptable to subtype classification of other viruses including dengue, influenza A, and hepatitis B and C virus.

bioinformatics

Multi-way methods for understanding longitudinal intervention effects on bacterial communities

BackgroundThis paper presents a strategy for statistical analysis and interpretation of longitudinal intervention effects on bacterial communities. Data from such experiments often suffers from small sample size, high degree of irrelevant variation, and missing data points. Our strategy is a combination of multi-way decomposition methods, multivariate ANOVA, multi-block regression, hierarchical clustering and phylogenetic network graphs. The aim is to provide answers to relevant research questions, which are both statistically valid and easy to interpret.\n\nResultsThe strategy is illustrated by analysing an intervention design where two mice groups were subjected to a treatment that caused inflammation in the intestines. Total microbiota in fecal samples was analysed at five time points, and the clinical end point was the load of colon cancer lesions. By using different combinations of the aforementioned methods, we were able to show that:\n\nO_LIThe treatment had a significant effect on the microbiota, and we have identified clusters of bacteria groups with different time trajectories.\nC_LIO_LIIndividual differences in the initial microbiota had a large effect on the load of tumors, but not on the formation of early-stage lesions (flat ACFs).\nC_LIO_LIThe treatment resulted in an increase in Bacteroidaceae, Prevotellaceae and Paraprevotellaceae, and this increase could be associated with the formation of cancer lesions.\nC_LI\n\nConclusionThe results show that by applying several data analytical methods in combination, we are able to view the system from different angles and thereby answer different research questions. We believe that multiway methods and multivariate ANOVA should be used more frequently in the bioinformatics fields, due to their ability to extract meaningful components from data sets with many collinear variables, few samples and a high degree of noise or irrelevant variation.

bioinformatics

Conserved SQ and QS motifs in bacterial effectors suggest pathogen interplay with the ATM kinase family during infection

Understanding how bacteria hijack eukaryotic cells during infection is vital to develop better strategies to counter the pathologies that they cause. ATM kinase family members phosphorylate eukaryotic protein substrates on Ser or Thr residues followed by Gln. The kinases are active under oxidative stress conditions and/or the presence of ds-DNA breaks. While examining the protein sequences of well-known bacterial effector proteins such as CagA and Tir, we noticed that they often show conserved (S/TQ) motifs, even though the evidence for effector phosphorylation by ATM has not been reported. We undertook a bioinformatics analysis to examine effectors for their potential to mimic the eukaryotic substrates of the ATM kinase. The candidates we found could interfere with the hosts intracellular signaling network upon interaction, which might give an advantage to the pathogen inside the host. Further, the putative phosphorylation sites should be accessible, conserved across species and, in the vicinity to the phosphorylation sites, positively charged residues should be depleted. We also noticed that the reverse motif (QT/S) is often also conserved and located close to (S/TQ) sites, indicating its potential biological role in ATM kinase function. Our findings could suggest a mechanism of infection whereby many pathogens inactivate/modulate the host ATM signaling pathway.

bioinformatics

Inferring the presence of aflatoxin-producing Aspergillus flavus strains using RNA sequencing and electronic probes as a transcriptomic screening tool

E-probe Diagnostic for Nucleic acid Analysis (EDNA) is a bioinformatic tool originally developed to detect plant pathogens in metagenomic databases. However, enhancements made to EDNA increased its capacity to conduct hypothesis directed detection of specific gene targets present in transcriptomic databases. To target specific pathogenicity factors used by the pathogen to infect its host or other targets of interest, e-probes need to be developed for transcripts related to that function. In this study, EDNA transcriptomics (EDNAtran) was developed to detect the expression of genes related to aflatoxin production at the transcriptomic level. E-probes were designed from genes up-regulated during A. flavus aflatoxin production. EDNAtran detected gene transcripts related to aflatoxin production in a transcriptomic database from corn, where aflatoxin was produced. The results were significantly different from e-probes being used in the transcriptomic database where aflatoxin was not produced (atoxigenic AF36 strain and toxigenic AF70 in Potato Dextrose Broth).

bioinformatics

Population assignment from cancer genome profiling data

For a variety of human malignancies, incidence, treatment efficacy and overall prognosis show considerable variation between different populations and ethnic groups. Disentangling the effects related to particular population backgrounds can help in both understanding cancer biology and in tailoring therapeutic interventions. Because self-reported or inferred patient data can be incomplete or misleading due to migration and genomic admixture, a data-driven ancestry estimation should be preferred. While algorithms to analyze ancestry structure from healthy individuals have been developed, an easy-to-use tool to assign population groups based on genotyping data from SNP profiles is still missing and benchmarking for the validity of population assignment strategy for aberrant cancer genomes was not tested.\n\nWe benchmarked the consistency and accuracy of cross-platform population assignment. We also demonstrated its high accuracy to process unaltered as well as cancer genomes. Despite widespread and extensive somatic mutations of cancer profiling data, population assignment consistency between germline and highly mutated samples from cancer patients reached of 97% and 92% for assignment into 5 and 26 populations re-spectively. Comparison of our benchmarked results with self-reported meta-data estimated a matching rate between 88% to 92%. Despite a relatively high matching rate, the ethnicity labels indicated in meta-data are vague compared to the standardized output from our tool.\n\nWe have developed a bioinformatics tool to assign the populations from genome profiling data and validated its performance in healthy as well as aberrant cancer genomes. It is ready-to-use for genotyping data from nine commercial SNP array platforms or sequencing data. This tool is effective to scrutinize the population structure in cancer genomes and provides better measure to integrate genotyping data from various platforms instead of self-reported information. It will facilitate research on interplay between ethnicity related genetic background and molecular patterns in cancer entities and disentangling possible hereditary contributions.\n\nThe docker image of the tool is provided in DockerHub as \"baudisgroup/snp2pop\".

bioinformatics

The impact of cigarette smoke exposure, and COPD or asthma status on ABC transporter gene expression in human airway epithelial cells

RationaleThe respiratory mucosa coordinates responses to infections, allergens, and exposures to air pollution. A relatively unexplored aspect of the respiratory mucosa are the expression and function of ATP Binding Cassette (ABC) transporters. ABC transporters are conserved in prokaryotes and eukaryotes, with humans expressing 48 transporters divided into 7 classes (ABCA, ABCB, ABCC, ABCD, ABDE, ABCF, and ABCG). Throughout the human body, ABC transporters regulate cAMP levels, chloride secretion, lipid transport, and anti-oxidant responses. A deeper exploration of the expression patterns of ABC transporters in the respiratory mucosa is warranted to determine their relevance in lung health and disease.\n\nMethodsWe used a bioinformatic approach complemented with in vitro experimental methods for validation of candidate ABC transporters. We analyzed the expression profiles of all 48 human ABC transporters in the respiratory mucosa using bronchial epithelial cell gene expression datasets available in NCBI GEO from well-characterized patient populations of healthy subjects and individuals that smoke cigarettes, or have been diagnosed with COPD or asthma. The Calu-3 airway epithelial cell line was used to interrogate selected results using a cigarette smoke extract exposure model.\n\nResultsUsing 9 distinct gene-expression datasets of primary human airway epithelial cells, we completed a focused analysis on 48 ABC transporters in samples from healthy subjects and individuals that smoke cigarettes, or have been diagnosed with COPD or asthma. In situ gene expression data demonstrate that ABC transporters are i) variably expressed in epithelial cells from different airway generations (top three expression levels - ABCA5, ABCA13, and ABCC5), ii) regulated by cigarette smoke exposure (ABCA13, ABCB6, ABCC1, and ABCC3), and iii) differentially expressed in individuals with COPD and asthma (ABCA13, ABCC1, ABCC2, ABCC9). An in vitro cell culture model of cigarette smoke exposure was able to recapitulate the in situ changes observed in cigarette smokers for ABCA13 and ABCC1.\n\nConclusionsOur in situ human gene expression data analysis reveals that ABC transporters are expressed throughout the airway generations in airway epithelial cells and can be modulated by environmental exposures important in chronic respiratory disease (e.g. cigarette smoking) and in individuals with chronic lung diseases (e.g. COPD or asthma). Our work highlights select ABC transporter candidates of interest and a relevant in vitro model that will enable a deeper understanding of the contribution of ABC transporters in the respiratory mucosa in lung health and disease.

bioinformatics

ParGenes: a tool for massively parallel model selection and phylogenetic tree inference on thousands of genes.

MotivationCoalescent- and reconciliation-based methods are now widely used to infer species phylogenies from genomic data. They typically use per-gene phylogenies as input, which requires conducting multiple individual tree inferences on a large set of multiple sequence alignments (MSAs). At present, no easy-to-use parallel tool for this task exists. Ad hoc scripts for this purpose do not only induce additional implementation overhead, but can also lead to poor resource utilization and long times-to-solution. We present ParGenes, a tool for simultaneously determining the best-fit model and inferring maximum likelihood (ML) phylogenies on thousands of independent MSAs using supercomputers.\n\nResultsParGenes executes common phylogenetic pipeline steps such as model-testing, ML inference(s), bootstrapping, and computation of branch support values via a single parallel program invocation. We evaluated ParGenes by inferring > 20, 000 phylogenetic gene trees with bootstrap support values from Ensembl Compara and VectorBase alignments in 28 hours on a cluster with 1024 nodes.\n\nAvailabilityGNU GPL at https://github.com/BenoitMorel/ParGenes.\n\nContactBenoit.Morel@h-its.org\n\nSupplementary informationSupplementary material is available at Bioinformatics online.

bioinformatics

PathwayMatcher: multi-omics pathway mapping and proteoform network generation

BackgroundMapping biomedical data to functional knowledge is an essential task in bioinformatics and can be achieved by querying identifiers, e.g. gene sets, in pathway knowledgebases. However, the isoform and post-translational modification states of proteins are lost when converting input and pathways into gene-centric lists.\n\nFindingsBased on the Reactome knowledgebase, we built a network of protein-protein interactions accounting for the documented isoform and modification statuses of proteins. We then implemented a command line application called PathwayMatcher (github.com/PathwayAnalysisPlatform/PathwayMatcher) to query this network. PathwayMatcher supports multiple types of omics data as input, and outputs the possibly affected biochemical reactions, subnetworks, and pathways.\n\nConclusionsPathwayMatcher enables refining the network-representation of pathways by including isoform and post-translational modifications. The specificity of pathway analyses is hence adapted to different levels of granularity and it becomes possible to distinguish interactions between different forms of the same protein.

bioinformatics

MOVIE: Multi-Omics VIsualization of Estimated contributions

SummaryThe growth of multi-omics datasets has given rise to many methods for identifying sources of common variation across data types. The unsupervised nature of these methods makes it difficult to evaluate their performance. We present MOVIE, Multi-Omics Visualization of Estimated contributions, as a framework for evaluating the degree of overfitting and the stability of unsupervised multi-omics methods. MOVIE plots the contributions of one data type against another to produce contribution plots, where contributions are calculated for each subject and each data type from the results of each multi-omics method. The usefulness of MOVIE is demonstrated by applying existing multi-omics methods to permuted null data and breast cancer data from The Cancer Genome Atlas. Contribution plots indicated that principal components-based Canonical Correlation Analysis overfit null data, while Sparse multiple Canonical Correlation Analysis and Multi-Omics Factor Analysis provided stable results with high specificity for both the real and permuted null datasets.\n\nAvailabilityMOVIE is available as an R package at https://github.com/mccabes292/movie\n\nContactmilove@email.unc.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics