Search bioRxivSearch

Biology subjects

Chan, T. E.

Publications and source records attributed to Chan, T. E..

4 recordsLinked to original sources

Signalling pathways drive heterogeneity of ground state pluripotency

Pluripotent stem cells (PSCs) can self-renew indefinitely while maintaining the ability to generate all cell types of the body. This plasticity is proposed to require heterogeneity in gene expression, driving a metastable state which may allow flexible cell fate choices. Contrary to this, naive PSC grown in fully defined 2i environmental conditions, containing small molecule inhibitors of MEK and GSK3 kinases, show homogenous pluripotency and lineage marker expression. However, here we show that 2i induces greater genome-wide heterogeneity than traditional serum-containing growth environments at the population level across both male and female PSCs. This heterogeneity is dynamic and reversible over time, consistent with a dynamic metastable equilibrium of the pluripotent state. We further show that the 2i environment causes increased heterogeneity in the calcium signalling pathway at both the population and single-cell level. Mechanistically, we identify loss of robustness regulators in the form of negative feedback to the upstream EGF receptor. Our findings advance the current understanding of the plastic nature of the pluripotent state and highlight the role of signalling pathways in the control of transcriptional heterogeneity. Furthermore, our results have critical implications for the current use of kinase inhibitors in the clinic, where inducing heterogeneity may increase the risk of cancer metastasis and drug resistance.

molecular biology

Empirical Bayes Meets Information Theoretical Network Reconstruction from Single Cell Data

Gene expression is controlled by networks of transcription factors and regulators, but the structure of these networks is as yet poorly understood and is thus inferred from data. Recent work has shown the efficacy of information theoretical approaches for network reconstruction from single cell transcriptomic data. Such methods use information to estimate dependence between every pair of genes in the dataset, then edges are inferred between top-scoring pairs. Dependence, however, does not indicate significance, and the definition of \"top-scoring\" is often arbitrary and a priori related to expected network size. This makes comparing networks across datasets difficult, because networks of a similar size are not necessarily similarly accurate. We present a method for performing formal hypothesis tests on putative network edges derived from information theory, bringing together empirical Bayes and work on theoretical null distributions for information measures. Thresholding based on empirical Bayes allows us to control network accuracy according to how we intend to use the network. Using single cell data from mouse pluripotent stem cells, we recover known interactions and suggest several new interactions for experimental validation (using a stringent threshold) and discover high-level interactions between sub-networks (using a more relaxed threshold). Furthermore, our method allows for the inclusion of prior information. We use in-silico data to show that even relatively poor quality prior information can increase the accuracy of a network, and demonstrate that the accuracy of networks inferred from single cell data can sometimes be improved by priors from population-level ChIP-Seq and qPCR data.

systems biology

Stem cell differentiation is a stochastic process with memory

Pluripotent stem cells are able to self-renew indefinitely in culture and differentiate into all somatic cell types in vivo. While much is known about the molecular basis of pluripotency, the molecular mechanisms of lineage commitment are complex and only partially understood. Here, using a combination of single cell profiling and mathematical modeling, we examine the differentiation dynamics of individual mouse embryonic stem cells (ESCs) as they progress from the ground state of pluripotency along the neuronal lineage. In accordance with previous reports we find that cells do not transit directly from the pluripotent state to the neuronal state, but rather first stochastically permeate an intermediate primed pluripotent state, similar to that found in the maturing epiblast in development. However, analysis of rate at which individual cells enter and exit this intermediate metastable state using a hidden Markov model reveals that the observed ESC and epiblast-like macrostates conceal a chain of unobserved cellular microstates, which individual cells transit through stochastically in sequence. These hidden microstates ensure that individual cells spend well-defined periods of time in each functional macrostate and encode a simple form of epigenetic memory that allows individual cells to record their position on the differentiation trajectory. To examine the generality of this model we also consider the differentiation of mouse hematopoietic stem cells along the myeloid lineage and observe remarkably similar dynamics, suggesting a general underlying process. Based upon these results we suggest a statistical mechanics view of cellular identities that distinguishes between functionally-distinct macrostates and the many functionally-similar molecular microstates associated with each macrostate. Taken together these results indicate that differentiation is a discrete stochastic process amenable to analysis using the tools of statistical mechanics.

systems biology

Network inference and hypotheses-generation from single-cell transcriptomic data using multivariate information measures

While single-cell gene expression experiments present new challenges for data processing, the cell-to-cell variability observed also reveals statistical relationships that can be used by information theory. Here, we use multivariate information theory to explore the statistical dependencies between triplets of genes in single-cell gene expression datasets. We develop PIDC, a fast, efficient algorithm that uses partial information decomposition (PID) to identify regulatory relationships between genes. We thoroughly evaluate the performance of our algorithm and demonstrate that the higher order information captured by PIDC allows it to outperform pairwise mutual information-based algorithms when recovering true relationships present in simulated data. We also infer gene regulatory networks from three experimental single-cell data sets and illustrate how network context, choices made during analysis, and sources of variability affect network inference. PIDC tutorials and open-source software for estimating PID are available here: https://github.com/Tchanders/network_inference_tutorials. PIDC should facilitate the identification of putative functional relationships and mechanistic hypotheses from single-cell transcriptomic data.

systems biology