Search bioRxivSearch

Biology subjects

Fabian J Theis

Publications and source records attributed to Fabian J Theis.

3 recordsLinked to original sources

cgCorrect: A method to correct for confounding cell-cell variation due to cell growth in single-cell transcriptomics

Motivation: Accessing gene expression at the single cell level has unraveled often large heterogeneity among seemingly homogeneous cells, which remained obscured in traditional population based approaches. The computational analysis of single-cell transcriptomics data, however, still imposes unresolved challenges with respect to normalization, visualization and modeling the data. One such issue are differences in cell size, which introduce additional variability into the data, for which appropriate normalization techniques are needed. Otherwise, these differences in cell size may obscure genuine heterogeneities among cell populations and lead to overdispersed steady-state distributions of mRNA transcript numbers.\n\nResults: We present cgCorrect, a statistical framework to correct for differences in cell size that are due to cell growth in single-cell transcriptomics data. We derive the probability for the cell growth corrected mRNA transcript number given the measured, cell size dependent mRNA transcript number, based on the assumption that the average number of transcripts in a cell increases proportional to the cells volume during cell cycle. cgCorrect can be used for both data normalization, and to analyze steady-state distributions used to infer the gene expression mechanism. We demonstrate its applicability on both simulated data and single-cell quantitative real-time PCR data from mouse blood stem and progenitor cells. We show that correcting for differences in cell size affects the interpretation of the data obtained by typically performed computational analysis.\n\nAvailability: A Matlab implementation of cgCorrect is available at http://icb.helmholtz-muenchen.de/cgCorrect\n\nSupplementary information: Supplementary information are available online. The simulated data set is available at http://icb.helmholtz-muenchen.de/cgCorrect

Bioinformatics

Exact Bayesian lineage tree-based inference identifies Nanog negative autoregulation in mouse embryonic stem cells

The autoregulatory motif of Nanog, a heterogeneously expressed core pluripotency factor in mouse embryonic stem cells, remains debated. Although recent time-lapse microscopy data provide the unparalleled ability to monitor Nanog expression at the single-cell level, the extraction of mechanistic knowledge is precluded by the lack of inference techniques suitable for noisy, incomplete and heterogeneous data obtained from proliferating cell populations.\n\nThis work identifies Nanogs autoregulatory motif from quantified time-lapse fluorescence line-age trees with STILT (Stochastic Inference on Lineage Trees), a novel particle-filter based algorithm for exact Bayesian parameter inference and model selection of stochastic models. We first verify STILTs ability to accurately infer parameters and select the correct autoregulatory motif from simulated data. We then apply STILT to time-lapse microscopy movies of a fluorescent Nanog fusion protein reporter and reject the possibility of positive autoregulation. Finally, we use STILT for experimental design, perform in silico overexpression simulations, and experimentally validate model predictions via exogenous Nanog overexpression. We finally conclude that the protein expression dynamics and overexpression experiments strongly suggest a weak negative feedback from the protein on the DNA activation rate.\n\nWe find that a simple autoregulatory mechanism can explain the observed heterogeneous Nanog dynamics. This finding has implications on the understanding of the core pluripotency network, such as supporting the ability of mESC populations to diversify their proteomic profile to respond to a spectrum of differentiation cues. Beyond this application STILT constitutes a generally applicable fully Bayesian approach for model selection of gene regulatory models on the basis of time-lapse imaging data of proliferating cell populations. STILT is freely available at: http://www.imsb.ethz.ch/research/claassen/Software/stilt--stochastic-inference-on-lineage-trees.html

Systems Biology

Network-based metabolite ratios for an improved functional characterization of genome-wide association study results

Genome-wide association studies (GWAS) with metabolite ratios as quantitative traits have successfully deepened our understanding of the complex relationship between genetic variants and metabolic phenotypes. Usually all ratio combinations are selected for association tests. However, with more metabolites being detectable, the quadratic increase of the ratio number becomes challenging from a statistical, computational and interpretational point-of-view. Therefore methods which select biologically meaningful ratios are required.\n\nWe here present a network-based approach by selecting only closely connected metabolites in a given metabolic network. The feasibility of this approach was tested on in silico data derived from simulated reaction networks. Especially for small effect sizes, network-based metabolite ratios (NBRs) improved the metabolite-based prediction accuracy of genetically-influenced reactions compared to the all ratios approach. Evaluating the NBR approach on published GWAS association results, we compared reported all ratio-SNP hits with results obtained by selecting only NBRs as candidates for association tests. Input networks for NBR selection were derived from public pathway databases or reconstructed from metabolomics data. NBR-candidates covered more than 80% of all significant ratio-SNP associations and we could replicate 7 out of 10 new associations predicted by the NBR approach.\n\nIn this study we evaluated a network-based approach to select biologically meaningful metabolite ratios as quantitative traits in GWAS. Taking metabolic network information into account facilitated the analysis and the biochemical interpretation of metabolite-gene association results. For upcoming studies, for instance with case-control design, large-scale metabolomics data and small sample numbers, the analysis of all possible metabolite ratios is not feasible due to the correction for multiple testing. Here our NBR approach increases the statistical power and lowers computational demands, allowing for a better understanding of the complex interplay between individual phenotypes, genetics and metabolic profiles.

Genetics