Search bioRxivSearch

Biology subjects

Ziegenhain, C.

Publications and source records attributed to Ziegenhain, C..

6 recordsLinked to original sources

TET1 drives global DNA demethylation via DPPA3-mediated inhibition of maintenance methylation

Genome-wide DNA demethylation is a unique feature of mammalian development and naive pluripotent stem cells. So far, it was unclear how mammals specifically achieve global DNA hypomethylation, given the high conservation of the DNA (de-)methylation machinery among vertebrates. We found that DNA demethylation requires TET activity but mostly occurs at sites where TET proteins are not bound suggesting a rather indirect mechanism. Among the few specific genes bound and activated by TET proteins was the naive pluripotency and germline marker Dppa3 (Pgc7, Stella), which undergoes TDG dependent demethylation. The requirement of TET proteins for genome-wide DNA demethylation could be bypassed by ectopic expression of Dppa3. We show that DPPA3 binds and displaces UHRF1 from chromatin and thereby prevents the recruitment and activation of the maintenance DNA methyltransferase DNMT1. We demonstrate that DPPA3 alone can drive global DNA demethylation when transferred to amphibians (Xenopus) and fish (medaka), both species that naturally do not have a Dppa3 gene and exhibit no post-fertilization DNA demethylation. Our results show that TET proteins are responsible for active and - indirectly also for - passive DNA demethylation; while TET proteins initiate local and gene-specific demethylation in vertebrates, the recent emergence of DPPA3 introduced a unique means of genome-wide passive demethylation in mammals and contributed to the evolution of epigenetic regulation during early mammalian development.

developmental biology

Strategies for quantitative RNA-seq analyses among closely related species

With the growing appreciation for the role of regulatory differences in evolution, researchers need to reliably quantify expression levels within and among species. However, for non-model organisms genome assemblies and annotations are often not available or have inferior quality, biasing the inference of expression changes to an unknown extent. Here, we explore the possibility to map RNA-seq reads from diverged species to one high quality reference genome. As test case, we used a small primate phylogeny ranging from Human to Marmoset spanning 12% nucleotide divergence. To distinguish the effect of sequence divergence and genome quality, we used in silico evolved genomes and existing genomes to simulate RNA-seq reads. These were then mapped to the genome of origin (self-mapping) as well as to one common reference (cross-mapping) to infer the quantification biases. We find that the bias due to cross-mapping is small for the closely related great apes ([≤] 4% divergence), and preferable to self-mapping given current genome qualities. For closely related species, cross-mapping provides easy access, high power and a well controlled false discovery rate for both; the analysis of intra-species expression differences as well as the detection of relative differences between species. If divergence increases, so that a substantial fraction of reads exceeds the limits of the mapper used, we find that gene-specific corrections and effect-size cutoffs can limit the bias before self-mapping becomes unavoidable. In summary, for the first time we systematically quantify biases in cross-species RNA-seq studies, providing guidance to best practices for these important evolutionary studies.

genomics

Virulence evolution in the opportunistic bacterial pathogen Pseudomonas aeruginosa

Bacterial opportunistic pathogens are feared for their difficult-to-treat nosocomial infections and for causing morbidity in immunocompromised patients. Here, we study how such a versatile opportunist, Pseudomonas aeruginosa, adapts to conditions inside and outside its model host Caenorhabditis elegans, and use phenotypic and genotypic screens to identify the mechanistic basis of virulence evolution. We found that virulence significantly dropped in unstructured environments both in the presence and absence of the host, but remained unchanged in spatially structured environments. Reduction of virulence was either driven by a substantial decline in the production of siderophores (in treatments without hosts) or toxins and proteases (in treatments with hosts). Whole-genome sequencing of evolved clones revealed positive selection and parallel evolution across replicates, and showed an accumulation of mutations in regulator genes controlling virulence factor expression. Our study identifies the spatial structure of the non-host environment as a key driver of virulence evolution in an opportunistic pathogen.

microbiology

mcSCRB-seq: sensitive and powerful single-cell RNA sequencing

Single-cell RNA sequencing (scRNA-seq) has emerged as the central genome-wide method to characterize cellular identities and processes. While performance of scRNA-seq methods is improving, an optimum in terms of sensitivity, cost-efficiency and flexibility has not yet been reached. Among the flexible plate-based methods \"Single-Cell RNA-Barcoding and Sequencing\" (SCRB-seq) is one of the most sensitive and efficient ones. Based on this protocol, we systematically evaluated experimental conditions such as reverse transcriptases, reaction enhancers and PCR polymerases. We find that adding polyethylene glycol considerably increases sensitivity by enhancing cDNA synthesis. Furthermore, using Terra polymerase increases efficiency due to a more even cDNA amplification that requires less sequencing of libraries. We combined these and other improvements to a new scRNA-seq library protocol we call \"molecular crowding SCRB-seq\" (mcSCRB-seq), which we show to be the most sensitive and one of the most efficient and flexible scRNA-seq methods to date.

genomics

zUMIs: A fast and flexible pipeline to process RNA sequencing data with UMIs

Single cell RNA-seq (scRNA-seq) experiments typically analyze hundreds or thousands of cells after amplification of the cDNA. The high throughput is made possible by the early introduction of sample-specific barcodes (BCs) and the amplification bias is alleviated by unique molecular identifiers (UMIs). Thus the ideal analysis pipeline for scRNA-seq data needs to efficiently tabulate reads according to both BC and UMI. zUMIs is such a pipeline, it can handle both known and random BCs and also efficiently collapses UMIs, either just for Exon mapping reads or for both Exon and Intron mapping reads. Another unique feature of zUMIs is the adaptive downsampling function, that facilitates dealing with hugely varying library sizes, but also allows to evaluate whether the library has been sequenced to saturation. zUMIs flexibility allows to accommodate data generated with any of the major scRNA-seq protocols that use BCs and UMIs. To illustrate the utility of zUMIs, we analysed a single-nucleus RNA-seq dataset and show that more than 35% of all reads map to Introns. We furthermore show that these intronic reads are informative about expression levels, significantly increasing the number of detected genes and improving the cluster resolution. Availability: https://github.com/sdparekh/zUMIs

bioinformatics

powsim: Power analysis for bulk and single cell RNA-seq experiments

Power analysis is essential to optimize the design of RNA-seq experiments and to assess and compare the power to detect differentially expressed genes in RNA-seq data. PowsimR is a flexible tool to simulate and evaluate differential expression from bulk and especially single-cell RNA-seq data making it suitable for a priori and posterior power analyses.

bioinformatics