Search bioRxivSearch

Biology subjects

Bar-Joseph, Z.

Publications and source records attributed to Bar-Joseph, Z..

9 recordsLinked to original sources

Dynamic interaction network inference from longitudinal microbiome data

BackgroundSeveral studies have focused on the microbiota living in environmental niches including human body sites. In many of these studies researchers collect longitudinal data with the goal of understanding not just the composition of the microbiome but also the interactions between the different taxa. However, analysis of such data is challenging and very few methods have been developed to reconstruct dynamic models from time series microbiome data.\n\nResultsHere we present a computational pipeline that enables the integration of data across individuals for the reconstruction of such models. Our pipeline starts by aligning the data collected for all individuals. The aligned profiles are then used to learn a dynamic Bayesian network which represents causal relationships between taxa and clinical variables. Testing our methods on three longitudinal microbiome data sets we show that our pipeline improve upon prior methods developed for this task. We also discuss the biological insights provided by the models which include several known and novel interactions.\n\nConclusionsWe propose a computational pipeline for analyzing longitudinal microbiome data. Our results provide evidence that microbiome alignments coupled with dynamic Bayesian networks improve predictive performance over previous methods and enhance our ability to infer biological relationships within the microbiome and between taxa and clinical factors.

microbiology

Cell lineage inference from SNP and scRNA-Seq data

Several recent studies focus on the inference of developmental and response trajectories from single cell NA-Seq (scRNA-Seq) data. A number of computational methods, often referred to as pseudo-time ordering, have been developed for this task. Recently, CRISPR has also been used to reconstruct lineage trees by inserting random mutations. However, both approaches suffer from drawbacks that limit their use. Here we develop a method to detect significant, cell type specific, sequence mutations from scRNA-Seq data. We show that only a few mutations are enough for reconstructing good branching models. Integrating these mutations with expression data further improves the accuracy of the reconstructed models. As we show, the majority of mutations we identify are likely RNA editing events indicating that such information can be used to distinguish cell types.

genomics

Continuous State HMMs for Modeling Time Series Single Cell RNA-Seq Data

MotivationMethods for reconstructing developmental trajectories from time series single cell RNA-Seq (scRNA-Seq) data can be largely divided into two categories. The first, often referred to as pseudotime ordering methods, are deterministic and rely on dimensionality reduction followed by an ordering step. The second learns a probabilistic branching model to represent the developmental process. While both types have been successful, each suffers from shortcomings that can impact their accuracy.\n\nResultsWe developed a new method based on continuous state HMMs (CSHMMs) for representing and modeling time series scRNA-Seq data. We define the CSHMM model and provide efficient learning and inference algorithms which allow the method to determine both the structure of the branching process and the assignment of cells to these branches. Analyzing several developmental single cell datasets we show that the CSHMM method accurately infers branching topology and correctly and continuously assign cells to paths, improving upon prior methods proposed for this task. Analysis of genes based on the continuous cell assignment identifies known and novel markers for different cell types.\n\nAvailabilitySoftware and Supporting website: www.andrew.cmu.edu/user/chiehll/CSHMM/\n\nContactzivbj@cs.cmu.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

CNNC: Convolutional Neural Networks for Co-Expression Analysis

Several methods were developed to mine gene-gene relationships from expression data. Examples include correlation and mutual information methods for co-expression analysis, clustering and undirected graphical models for functional assignments and directed graphical models for pathway reconstruction. Using a novel encoding for gene expression data, followed by deep neural networks analysis, we present a framework that can successfully address all these diverse tasks. We show that our method, CNNC, improves upon prior methods in tasks ranging from predicting transcription factor targets to identifying disease related genes to causality inference. CNNCs encoding provides insights about some of the decisions it makes and their biological basis. CNNC is flexible and can easily be extended to integrate additional types of genomics data leading to further improvements in its performance.\n\nSupporting website with software and data: https://github.com/xiaoyeye/CNNC.

systems biology

scQuery: a web server for comparative analysis of single-cell RNA-seq data

Single cell RNA-Seq (scRNA-seq) studies often profile upward of thousands of cells in heterogeneous environments. Current methods for characterizing cells perform unsupervised analysis followed by assignment using a small set of known marker genes. Such approaches are limited to a few, well characterized cell types. To enable large scale supervised characterization we developed an automated pipeline to download, process, and annotate publicly available scRNA-seq datasets. We extended supervised neural networks to obtain efficient and accurate representations for scRNA-seq data. We applied our pipeline to analyze data from over 500 different studies with over 300 unique cell types and show that supervised methods greatly outperform unsupervised methods for cell type identification. A case study of neural degeneration data highlights the ability of these methods to identify differences between cell type distributions in healthy and diseased mice. We implemented a web server that compares new datasets to collected data employing fast matching methods in order to determine cell types, key genes, similar prior studies, and more.

bioinformatics

Proteome-scale detection of drug-target interactions using correlations in transcriptomic perturbations

The development of an expanded chemical space for screening is an essential step in the challenge of identifying chemical probes for new, genomic-era protein targets. However, the difficulty of identifying targets for novel compounds leads to the prioritization of synthesis linked to known active scaffolds that bind familiar protein families, slowing the exploration of available chemical space. To change this paradigm, we validated a new pipeline capable of identifying compound-protein interactions even for compounds with no similarity to known drugs. Based on differential mRNA profiles from drug treatments and gene knockdowns across multiple cell types, we show that drugs cause gene regulatory network effects that correlate with those produced by silencing their target protein-coding gene. Applying supervised machine learning to exploit compound-knockdown signature correlations and enriching our predictions using an orthogonal structure-based screen, we achieved top-10/top-100 target prediction accuracies of 26%/41%, respectively, on a validation set 152 FDA-approved drugs and 3104 potential targets. We further predicted targets for 1680 compounds and validated a total of seven novel interactions with four difficult targets, including non-covalent modulators of HRAS and KRAS. We found that drug-target interactions manifest as gene expression correlations between drug treatment and both target gene knockdown and up/down-stream knockdowns. These correlations provide biologically relevant insight on the cell-level impact of disrupting protein interactions, highlighting the complex genetic phenotypes of drug treatments. Our pipeline can accelerate the identification and development of novel chemistries with potential to become drugs by screening for compound-target interactions in the full human interactome.

genomics

Hercules: a profile HMM-based hybrid error correction algorithm for long reads

MotivationChoosing whether to use second or third generation sequencing platforms can lead to trade-offs between accuracy and read length. Several studies require long and accurate reads including de novo assembly, fusion and structural variation detection. In such cases researchers often combine both technologies and the more erroneous long reads are corrected using the short reads. Current approaches rely on various graph based alignment techniques and do not take the error profile of the underlying technology into account. Memory- and time-efficient machine learning algorithms that address these shortcomings have the potential to achieve better and more accurate integration of these two technologies.\n\nResultsWe designed and developed Hercules, the first machine learning-based long read error correction algorithm. The algorithm models every long read as a profile Hidden Markov Model with respect to the underlying platforms error profile. The algorithm learns a posterior transition/emission probability distribution for each long read and uses this to correct errors in these reads. Using datasets from two DNA-seq BAC clones (CH17-157L1 and CH17-227A2), and human brain cerebellum polyA RNA-seq, we show that Hercules-corrected reads have the highest mapping rate among all competing algorithms and highest accuracy when most of the basepairs of a long read are covered with short reads.\n\nAvailabilityHercules source code is available at https://github.com/BilkentCompGen/Hercules

bioinformatics

Cardiac directed differentiation using small molecule Wnt modulation at single-cell resolution

Differentiation into diverse cell lineages requires the orchestration of gene regulatory networks guiding diverse cell fate choices. Utilizing human pluripotent stem cells, we measured expression dynamics of 17,718 genes from 43,168 cells across five time points over a thirty day time-course of in vitro cardiac-directed differentiation. Unsupervised clustering and lineage prediction algorithms were used to map fate choices and transcriptional networks underlying cardiac differentiation. We leveraged this resource to identify strategies for controlling in vitro differentiation as it occurs in vivo. HOPX, a non-DNA binding homeodomain protein essential for heart development in vivo was identified as dys-regulated in in vitro derived cardiomyocytes. Utilizing genetic gain and loss of function approaches, we dissect the transcriptional complexity of the HOPX locus and identify the requirement of hypertrophic signaling for HOPX transcription in hPSC-derived cardiomyocytes. This work provides a single cell dissection of the transcriptional landscape of cardiac differentiation for broad applications of stem cells in cardiovascular biology.

developmental biology

Using Neural Networks To Improve Single-Cell RNA-Seq Data Analysis

While only recently developed, the ability to profile expression data in single cells (scRNA-Seq) has already led to several important studies and findings. However, this technology has also raised several new computational challenges including questions related to handling the noisy and sometimes incomplete data, how to identify unique group of cells in such experiments and how to determine the state or function of specific cells based on their expression profile. To address these issues we develop and test a method based on neural networks (NN) for the analysis and retrieval of single cell RNA-Seq data. We tested various NN architectures, some biologically motivated, and used these to obtain a reduced dimension representation of the single cell expression data. We show that the NN method improves upon prior methods in both, the ability to correctly group cells in experiments not used in the training and the ability to correctly infer cell type or state by querying a database of tens of thousands of single cell profiles. Such database queries (which can be performed using our web server) will enable researchers to better characterize cells when analyzing heterogeneous scRNA-Seq samples.\n\nSupporting website: http://sb.cs.cmu.edu/scnn/\n\nPassword for accessing the retrieval task webserver: scRNA-Seq

bioinformatics