Search bioRxivSearch

Biology subjects

Gut, I.

Publications and source records attributed to Gut, I..

7 recordsLinked to original sources

Inference of genomic spatial organization from a whole genome bisulfite sequencing sample.

Common approaches to characterize the structure of the DNA in the nucleus, such as the different Chromosome Conformation Capture methods, have not currently been widely applied to different tissue types due to several practical difficulties including the requirement for intact cells to start the sample preparation. In contrast, techniques based on sodium bisulfite conversion of DNA to assay DNA methylation, have been widely applied to many different tissue types in a variety of organisms. Recent work has shown the possibility of inferring some aspects of the three dimensional DNA structure from DNA methylation data, raising the possibility of three dimensional DNA structure prediction using the large collection of already generated DNA methylation datasets. We propose a simple method to predict the values of the first eigenvector of the Hi-C matrix of a sample (and hence the positions of the A and B compartments) using only the GC content of the sequence and a single whole genome bisulfite sequencing (WGBS) experiment which yields information on the methylation levels and their variability along the genome. We train and test our model on 10 samples for which we have data from both bisulfite sequencing and chromosome conformation experiments and our most relevant finding is that the variability of DNA methylation along the sequence is often a better predictor than methylation itself. We then run a prediction on 206 DNA methylation profiles produced by the Blueprint project and use ChIP-Seq and RNA-Seq data to confirm that the forecasted eigenvector delineates correctly the physical chromatin compartments observed with the Hi-C experiment.

bioinformatics

Single cell expression analysis uncouples transdifferentiation and reprogramming

Many somatic cell types are plastic, having the capacity to convert into other specialized cells (transdifferentiation)(1) or into induced pluripotent stem cells (iPSCs, reprogramming)(2) in response to transcription factor over-expression. To explore what makes a cell plastic and whether these different cell conversion processes are coupled, we exposed bone marrow derived pre-B cells to two different transcription factor overexpression protocols that efficiently convert them either into macrophages or iPSCs and monitored the two processes over time using single cell gene expression analysis. We found that even in these highly efficient cell fate conversion systems, cells differ in both their speed and path of transdifferentiation and reprogramming. This heterogeneity originatesin two starting pre-B cell subpopulations,large pre-BII and the small pre-BII cells they normally differentiate into. The large cells transdifferentiate slowly but exhibit a high efficiency of iPSC reprogramming. In contrast, the small cells transdifferentiate rapidly but are highly resistant to reprogramming. Moreover, the large B cells induce a stronger transient granulocyte/macrophage progenitor (GMP)-like state, while the small B cells undergo a more direct conversion to the macrophage fate. The large cells are cycling and exhibit high Myc activity whereas the small cells are Myc low and mostly quiescent. The observed heterogeneity of the two cell conversion processes can therefore be traced to two closely related cell types in the starting population that exhibit different types of plasticity. These data show that a somatic cells propensity for either transdifferentiation and reprogramming can be uncoupled.\n\nOne sentence summarySingle cell transcriptomics of cell conversions

developmental biology

Selective single molecule sequencing and assembly of a human Y chromosome of African origin

Mammalian Y chromosomes are often neglected from genomic analysis. Due to their inherent assembly difficulties, high repeat content, and large ampliconic regions1, only a handful of species have their Y chromosome properly characterized. To date, just a single human reference quality Y chromosome, of European ancestry, is available due to a lack of accessible methodology2-5. To facilitate the assembly of such complicated genomic territory, we developed a novel strategy to sequence native, unamplified flow sorted DNA on a MinION nanopore sequencing device. Our approach yields a highly continuous and complete assembly of the first human Y chromosome of African origin. It constitutes a significant improvement over comparable previous methods, increasing continuity by more than 800%6, thus allowing a chromosome scale analysis of human Y chromosomes. Sequencing native DNA also allows to take advantage of the nanopore signal data to detect epigenetic modifications in situ7. This approach is in theory generalizable to any species simplifying the assembly of extremely large and repetitive genomes.

genomics

matchSCore: Matching Single-Cell Phenotypes Across Tools and Experiments

Single-cell transcriptomics allows the identification of cellular types, subtypes and states through cell clustering. In this process, similar cells are grouped before determining co-expressed marker genes for phenotype inference. The performance of computational tools is directly associated to their marker identification accuracy, but the lack of an optimal solution challenges a systematic method comparison. Moreover, phenotypes from different studies are challenging to integrate, due to varying resolution, methodology and experimental design. In this work we introduce matchSCore (https://github.com/elimereu/matchSCore), an approach to match cell populations fast across tools, experiments and technologies. We compared 14 computational methods and evaluated their accuracy in clustering and gene marker identification in simulated data sets. We further used matchSCore to project cell type identities across mouse and human cell atlas projects. Despite originating from different technologies, cell populations could be matched across data sets, allowing the assignment of clusters to reference maps and their annotation.

bioinformatics

Epigenomic and functional dynamics of human bone marrow myeloid differentiation to mature blood neutrophils

Neutrophils are short-lived blood cells that play a critical role in host defense against infections. To better comprehend neutrophil functions and their regulation, we provide a complete epigenetic and functional overview of their differentiation stages from bone marrow-residing progenitors to mature circulating cells. Integration of epigenetic and transcriptome dynamics reveals an enforced regulation of differentiation, through cellular functions such as: release of proteases, respiratory burst, cell cycle regulation and apoptosis. We observe an early establishment of the cytotoxic capability, whilst the signaling components that activate antimicrobial mechanisms are transcribed at later stages, outside the bone marrow, thus preventing toxic effects in the bone marrow niche. Altogether, these data reveal how the developmental dynamics of the epigenetic landscape orchestrate the daily production of large number of neutrophils required for innate host defense and provide a comprehensive overview of the epigenomes of differentiating human neutrophils.\n\nKey pointsO_LIDynamic acetylation enforces human neutrophil progenitor differentiation.\nC_LI\n\nO_LINeutrophils cytotoxic capability is established early at the (pro)myelocyte stage.\nC_LI\n\nO_LICoordinated signaling component expression prevents unwanted toxic effects to the bone marrow niche.\nC_LI

immunology

DNA methylation oscillation defines classes of enhancers

Understanding the regulatory landscape of human cells requires the integration of genomic and epigenomic maps, capturing combinatorial levels of cell type-specific and invariant activity states.\n\nHere, we segmented whole-genome bisulfite sequencing-derived methylomes into consecutive blocks of co-methylation (COMETs) to obtain spatial variation patterns of DNA methylation (DNAm oscillations) integrated with histone modifications and promoter-enhancer interactions derived from promoter capture Hi-C (PCHi-C) sequencing of the same purified blood cells.\n\nMapping DNAm oscillations onto regulatory genome annotation revealed that enhancers are enriched for DNAm hyper-oscillations (>30-fold), where multiple machine learning models support DNAm as predictive of enhancer location. Based on this analysis, we report overall predictive power of 99% for DNAm oscillations, 77.3% for DNaseI, 41% for CGIs, 20% for UMRs and 0% for LMRs, demonstrating the power of DNAm oscillations over other methods for enhancer prediction. Methylomes of activated and non-activated CD4+ T cells indicate that DNAm oscillations exist in both states irrespective of activation; hence they can be used to determine the location of latent enhancers.\n\nOur approach advances the identification of tissue-specific regulatory elements and outperforms previous approaches defining enhancer classes based on DNA methylation.

genomics

bigSCale: An Analytical Framework for Big-Scale Single-Cell Data

Single-cell RNA sequencing significantly deepened our insights into complex tissues and latest techniques are capable processing ten-thousands of cells simultaneously. With bigSCale, we provide an analytical framework being scalable to analyze millions of cells, addressing challenges of future large datasets. Unlike previous methods, bigSCale does not constrain data to fit an a priori-defined distribution and instead uses an accurate numerical model of noise. We evaluated the performance of bigSCale using a biological model of aberrant gene expression in patient derived neuronal progenitor cells and simulated datasets, which underlined its speed and accuracy in differential expression analysis. We further applied bigSCale to analyze 1.3 million cells from the mouse developing forebrain. Herein, we identified rare populations, such as Reelin positive Cajal-Retzius neurons, for which we determined a previously not recognized heterogeneity associated to distinct differentiation stages, spatial organization and cellular function. Together, bigSCale presents a perfect solution to address future challenges of large single-cell datasets.\n\nExtended AbstractSingle-cell RNA sequencing (scRNAseq) significantly deepened our insights into complex tissues by providing high-resolution phenotypes for individual cells. Recent microfluidic-based methods are scalable to ten-thousands of cells, enabling an unbiased sampling and comprehensive characterization without prior knowledge. Increasing cell numbers, however, generates extremely big datasets, which extends processing time and challenges computing resources. Current scRNAseq analysis tools are not designed to analyze datasets larger than from thousands of cells and often lack sensitivity and specificity to identify marker genes for cell populations or experimental conditions. With bigSCale, we provide an analytical framework for the sensitive detection of population markers and differentially expressed genes, being scalable to analyze millions of single cells. Unlike other methods that use simple or mixture probabilistic models with negative binomial, gamma or Poisson distributions to handle the noise and sparsity of scRNAseq data, bigSCale does not constrain the data to fit an a priori-defined distribution. Instead, bigSCale uses large sample sizes to estimate a highly accurate and comprehensive numerical model of noise and gene expression. The framework further includes modules for differential expression (DE) analysis, cell clustering and population marker identification. Moreover, a directed convolution strategy allows processing of extremely large data sets, while preserving the transcript information from individual cells.\n\nWe evaluate the performance of bigSCale using a biological model for reduced or elevated gene expression levels. Specifically, we perform scRNAseq of 1,920 patient derived neuronal progenitor cells from Williams-Beuren and 7q11.23 microduplication syndrome patients, harboring a deletion or duplication of 7q11.23, respectively. The affected region contains 28 genes whose transcriptional levels vary in line with their allele frequency. BigSCale detects expression changes with respect to cells from a healthy donor and outperforms other methods for single-cell DE analysis in sensitivity. Simulated data sets, underline the performance of bigSCale in DE analysis as it is faster and more sensitive and specific than other methods. The probabilistic model of cell-distances within bigSCale is further suitable for unsupervised clustering and the identification of cell types and subpopulations. Using bigSCale, we identify all major cell types of the somatosensory cortex and hippocampus analyzing 3,005 cells from adult mouse brains. Remarkably, we increase the number of cell population specific marker genes 4-6-fold compared to the original analysis and, moreover, define markers of higher order cell types. These include CD90 (Thy1), a neuronal surface receptor, potentially suitable for isolating intact neurons from complex brain samples.\n\nTo test its applicability for large data sets, we apply bigSCale on scRNAseq data from 1.3 million cells derived from the pallium of the mouse developing forebrain (E18, 10x Genomics). Our directed down-sampling strategy accumulates transcript counts from cells with similar transcriptional profiles into index cell transcriptomes, thereby defining cellular clusters with improved resolution. Accordingly, index cell clusters provide a rich resource of marker genes for the main brain cell types and less frequent subpopulations. Our analysis of rare populations includes poorly characterized developmental cell types, such as neuron progenitors from the subventricular zone and neocortical Reelin positive neurons known as Cajal-Retzius (CR) cells. The latter represent a transient population which regulates the laminar formation of the developing neocortex and whose malfunctioning causes major neurodevelopmental disorders like autism or schizophrenia. Most importantly, index cell cluster can be deconvoluted to individual cell level for targeted analysis of populations of interest. Through decomposition of Reelin positive neurons, we determined a previously not recognized heterogeneity among CR cells, which we could associate to distinct differentiation stages as well as spatial and functional differences in the developing mouse brain. Specifically, subtypes of CR cells identified by bigSCale express different compositions of NMDA, AMPA and glycine receptor subunits, pointing to subpopulations with distinct membrane properties. Furthermore, we found Cxcl12, a chemokine secreted by the meninges and regulating the tangential migration of CR cells, to be also expressed in CR cells located in the marginal zone of the neocortex, indicating a self-regulated migration capacity.\n\nTogether, bigSCale presents a perfect solution for the processing and analysis of scRNAseq data from millions of single cells. Its speed and sensitivity makes it suitable to the address future challenges of large single-cell data sets.

genomics