Search bioRxivSearch

Biology subjects

Mold, J. E.

Publications and source records attributed to Mold, J. E..

2 recordsLinked to original sources

Conbase: a software for discovery of clonal somatic mutations in single cells through read phasing

Here we report the development of Conbase, a software application for the identification of somatic mutations in single cell DNA sequencing data with high rates of allelic dropout and at low read depth. Conbase leverages data from multiple samples in a dataset and utilizes read phasing to call somatic single nucleotide variants and to accurately predict genotypes in whole genome amplified single cells in somatic variant loci. We demonstrate the accuracy of Conbase on simulated datasets, in vitro expanded fibroblasts and clonally in vivo expanded lymphocyte populations isolated directly from a healthy human donor.

bioinformatics

Probabilistic Count Matrix Factorization for Single Cell Expression Data Analysis

The development of high throughput single-cell technologies now allows the investigation of the genome-wide diversity of transcription. This diversity has shown two faces: the expression dynamics (gene to gene variability) can be quantified more accurately, thanks to the measurement of lowly-expressed genes. Second, the cell-to-cell variability is high, with a low proportion of cells expressing the same gene at the same time/level. Those emerging patterns appear to be very challenging from the statistical point of view, especially to represent and to provide a summarized view of single-cell expression data. PCA is one of the most powerful frameworks to provide a suitable representation of high dimensional datasets, by searching for new axes catching the most variability in the data. Unfortunately, classical PCA is based on Euclidean distances and projections that work poorly in presence of over-dispersed counts showing zero-inflation. We propose a probabilistic Count Matrix Factorization (pCMF) approach for single-cell expression data analysis, that relies on a sparse Gamma-Poisson factor model. This hierarchical model is inferred using a variational EM algorithm. We show how this probabilistic framework induces a geometry that is suitable for single-cell data, and produces a compression of the data that is very powerful for clustering purposes. Our method is competed to other standard representation methods like ZIFA and t-SNE, and we illustrate its performance on simulated and publicly available data.

bioinformatics