Search bioRxivSearch

Biology subjects

Tegner, J.

Publications and source records attributed to Tegner, J..

4 recordsLinked to original sources

Building gene regulatory networks from single-cell ATAC-seq and RNA-seq using Linked Self-Organizing Maps

Rapid advances in single-cell assays have outpaced methods for analysis of those data types. Different single-cell assays show extensive variation in sensitivity and signal to noise levels. In particular, scATAC-seq generates extremely sparse and noisy datasets. Existing methods developed to analyze this data require cells amenable to pseudo-time analysis or require datasets with drastically different cell-types. We describe a novel approach using self-organizing maps (SOM) to link scATAC-seq and scRNA-seq data that overcomes these challenges and can generate draft regulatory networks. Our SOMatic package generates chromatin and gene expression SOMs separately and combines them using a linking function. We applied SOMatic on a mouse pre-B cell differentiation time-course using controlled Ikaros over-expression to recover gene ontology enrichments, identify motifs in genomic regions showing similar single-cell profiles, and generate a gene regulatory network that both recovers known interactions and predicts new Ikaros targets during the differentiation process. The ability of linked SOMs to detect emergent properties from multiple types of highly-dimensional genomic data with very different signal properties opens new avenues for integrative analysis of single-cells.

genomics

The ultra-sensitive Nodewalk technique identifies stochastic from virtual, population-based enhancer hubs regulating MYC in 3D: Implications for the fitness of cancer cells

The relationship between stochastic transcriptional bursts and dynamic 3D chromatin states is not well understood due to poor sensitivity and/or resolution of current chromatin structure-based assays. Consequently, it is not well established if enhancers operate individually and/or in clusters to coordinate gene transcription. In the current study, we introduce Nodewalk, which uniquely combines high sensitivity with high resolution to enable the analysis of chromatin networks in minute input material. The >10,000-fold increase in sensitivity over other many-to-all competing methods uncovered that active chromatin hubs identified in large input material, corresponding to 10 000 cells, flanking the MYC locus are primarily virtual. Thus, the close agreement between chromatin interactomes generated from aliquots corresponding to less than 10 cells with randomly re-sampled interactomes, we find that numerous distal enhancers positioned within flanking topologically associating domains (TADs) converge on MYC in largely mutually exclusive manners. Moreover, when comparing with several enhancer baits, the assignment of the MYC locus as the node with the highest dynamic importance index, indicates that it is MYC targeting its enhancers, rather than vice versa. Dynamic changes in the configuration of the boundary between TADs flanking MYC underlie numerous stochastic encounters with a diverse set of enhancers to depict the plasticity of its transcriptional regulation. Such an arrangement might increase the fitness of the cancer cell by increasing the probability of MYC transcription in response to a wide range of environmental cues encountered by the cell during the neoplastic process.

genomics

An Algorithmic Information Calculus for Causal Discovery and Reprogramming Systems

We introduce a new conceptual framework and a model-based interventional calculus to steer, manipulate, and reconstruct the dynamics and generating mechanisms of non-linear dynamical systems from partial and disordered observations based on the contributions of each of the systems, by exploiting first principles from the theory of computability and algorithmic information. This calculus entails finding and applying controlled interventions to an evolving object to estimate how its algorithmic information content is affected in terms of positive or negative shifts towards and away from randomness in connection to causation. The approach is an alternative to statistical approaches for inferring causal relationships and formulating theoretical expectations from perturbation analysis. We find that the algorithmic information landscape of a system runs parallel to its dynamic attractor landscape, affording an avenue for moving systems on one plane so they can be controlled on the other plane. Based on these methods, we advance tools for reprogramming a system that do not require full knowledge or access to the systems actual kinetic equations or to probability distributions. This new approach yields a suite of universal parameter-free algorithms of wide applicability, ranging from the discovery of causality, dimension reduction, feature selection, model generation, a maximal algorithmic-randomness principle and a systems (re)programmability index. We apply these methods to static (e.coli Transcription Factor network) and to evolving genetic regulatory networks (differentiating naive from Th17 cells, and the CellNet database). We highlight their ability to pinpoint key elements (genes) related to cell function and cell development, conforming to biological knowledge from experimentally validated data and the literature, and demonstrate how the method can reshape a systems dynamics in a controlled manner through algorithmic causal mechanisms.

systems biology

Predicting Causal Relationships from Biological Data: Applying Automated Casual Discovery on Mass Cytometry Data of Human Immune Cells

Learning the causal relationships that define a molecular system allows us to predict how the system will respond to different interventions. Distinguishing causality from mere association typically requires randomized experiments. Methods for automated causal discovery from limited experiments exist, but have so far rarely been tested in systems biology applications. In this work, we apply state-of-the art causal discovery methods on a large collection of public mass cytometry data sets, measuring intra-cellular signaling proteins of the human immune system and their response to several perturbations. We show how different experimental conditions can be used to facilitate causal discovery, and apply two fundamental methods that produce context-specific causal predictions. Causal predictions were reproducible across independent data sets from two different studies, but often disagree with the KEGG pathway databases. Within this context, we discuss the caveats we need to overcome for automated causal discovery to become a part of the routine data analysis in systems biology.

bioinformatics