Search bioRxivSearch

Biology subjects

Dermitzakis, E. T.

Publications and source records attributed to Dermitzakis, E. T..

6 recordsLinked to original sources

Single-cell transcriptomics of the mouse gonadal soma reveals the establishment of sexual dimorphism in distinct cell lineages

Sex determination is a unique process that allows the study of multipotent progenitors and their acquisition of sex-specific fates during differentiation of the gonad into a testis or an ovary. Using time-series single-cell RNA sequencing (scRNA-seq) on ovarian Nr5a1-GFP+ somatic cells during sex determination, we identified a single population of early progenitors giving rise to both pre-granulosa cells and potential steroidogenic precursor cells. By comparing time-series scRNA-seq of XX and XY somatic cells, we demonstrate that the supporting cells emerge from the early progenitors with a non-sex-specific transcriptomic program, before pre-granulosa and Sertoli cells acquire their sex-specific identity. In XX and XY steroidogenic precursors similar transcriptomic profiles underlie the acquisition of cell fate, but with a delay in XX cells. Our data provide a novel framework, at single-cell resolution, for further interrogation of the molecular and cellular basis of mammalian sex determination.

developmental biology

Expression estimation and eQTL mapping for HLA genes with a personalized pipeline

The HLA (Human Leukocyte Antigens) genes are well-documented targets of balancing selection, and variation at these loci is associated with many disease phenotypes. Variation in expression levels also influences disease susceptibility and resistance, but little information exists about the regulation and population-level patterns of expression due to the difficulty in mapping short reads to these highly polymorphic loci, and in accounting for the existence of several paralogues. We developed a computational pipeline to accurately estimate expression for HLA genes based on RNA-seq, improving both locus-level and allele-level estimates. First, reads are aligned to all known HLA sequences in order to infer HLA genotypes, then quantification of expression is carried out using a personalized index. We use simulations to show that expression estimates are not biased due to divergence from the reference genome. We applied our pipeline to GEUVADIS dataset, and compared the quantifications to those obtained with reference transcriptome, and found that a substantial portion of the variation captured by the HLA-personalized index in not captured by the standard index (23%). We describe the impact of the HLA-personalized approach on downstream analyses for seven HLA loci (HLA-A, HLA-B, HLA-C, HLA-DPB1, HLA-DQA1, HLA-DQB1, HLA-DRB1). Although the influence of the HLA-personalized approach is modest for eQTL mapping, the p-values and the causality of the eQTLs obtained are better than when the reference transcriptome is used. Finally, we integrate information on HLA-allele level expression with the eQTL findings to show that the HLA allele is an important layer of variation to understand HLA regulation.

genomics

Discovery of biomarkers for glycaemic deterioration before and after the onset of type 2 diabetes: an overview of the data from the epidemiological studies within the IMI DIRECT Consortium

Abstract/SummaryO_ST_ABSBackground and aimsC_ST_ABSUnderstanding the aetiology, clinical presentation and prognosis of type 2 diabetes (T2D) and optimizing its treatment might be facilitated by biomarkers that help predict a persons susceptibility to the risk factors that cause diabetes or its complications, or response to treatment. The IMI DIRECT (Diabetes Research on Patient Stratification) Study is a European Union (EU) Innovative Medicines Initiative (IMI) project that seeks to test these hypotheses in two recently established epidemiological cohorts. Here, we describe the characteristics of these cohorts at baseline and at the first main follow-up examination (18-months).\n\nMaterials and methodsFrom a sampling-frame of 24,682 European-ancestry adults in whom detailed health information was available, participants at varying risk of glycaemic deterioration were identified using a risk prediction algorithm and enrolled into a prospective cohort study (n=2127) undertaken at four study centres across Europe (Cohort 1: prediabetes). We also recruited people from clinical registries with recently diagnosed T2D (n=789) into a second cohort study (Cohort 2: diabetes). The two cohorts were studied in parallel with matched protocols. Endogenous insulin secretion and insulin sensitivity were modelled from frequently sampled 75g oral glucose tolerance (OGTT) in Cohort 1 and with mixed-meal tolerance tests (MMTT) in Cohort 2. Additional metabolic biochemistry was determined using blood samples taken when fasted and during the tolerance tests. Body composition was assessed using MRI and lifestyle measures through self-report and objective methods.\n\nResultsUsing ADA-2011 glycaemic categories, 33% (n=693) of Cohort 1 (prediabetes) had normal glucose regulation (NGR), and 67% (n=1419) had impaired glucose regulation (IGR). 76% of the cohort was male, age=62(6.2) years; BMI=27.9(4.0) kg/m2; fasting glucose=5.7(0.6) mmol/l; 2-hr glucose=5.9(1.6) mmol/l [mean(SD)]. At follow-up, 18.6(1.4) months after baseline, fasting glucose=5.8(0.6) mmol/l; 2-hr OGTT glucose=6.1(1.7) mmol/l [mean(SD)]. In Cohort 2 (diabetes): 65% (n=508) were lifestyle treated (LS) and 35% (n=271) were lifestyle + metformin treated (LS+MET). 58% of the cohort was male, age=62(8.1) years; BMI=30.5(5.0) kg/m2; fasting glucose=7.2(1.4)mmol/l; 2-hr glucose=8.6(2.8) mmol/l [mean(SD)]. At follow-up, 18.2(0.6) months after baseline, fasting glucose=7.8(1.8) mmol/l; 2-hr MMTT glucose=9.5(3.3) mmol/l [mean(SD)].\n\nConclusionThe epidemiological IMI DIRECT cohorts are the most intensely characterised prospective studies of glycaemic deterioration to date. Data from these cohorts help illustrate the heterogeneous characteristics of people at risk of or with T2D, highlighting the rationale for biomarker stratification of the disease - the primary objective of the IMI DIRECT consortium.\n\nAbbreviations

epidemiology

Genomic dissection of Systemic Lupus Erythematosus: Distinct Susceptibility, Activity and Severity Signatures

Recent genetic and genomics approaches have yielded novel insights in the pathogenesis of Systemic Lupus Erythematosus (SLE) but the diagnosis, monitoring and treatment still remain largely empirical1,2. We reasoned that molecular characterization of SLE by whole blood transcriptomics may facilitate early diagnosis and personalized therapy. To this end, we analyzed genotypes and RNA-seq in 142 patients and 58 matched healthy individuals to define the global transcriptional signature of SLE. By controlling for the estimated proportions of circulating immune cell types, we show that the Interferon (IFN) and p53 pathways are robustly expressed. We also report cell-specific, disease-dependent regulation of gene expression and define a core/susceptibility and a flare/activity disease expression signature, with oxidative phosphorylation, ribosome regulation and cell cycle pathways being enriched in lupus flares. Using these data, we define a novel index of disease activity/severity by combining the validated Systemic Lupus Erythematosus Disease Activity Index (SLEDAI)1 with a new variable derived from principal component analysis (PCA) of RNA-seq data. We also delineate unique signatures across disease endo-phenotypes whereby active nephritis exhibits the most extensive changes in transcriptome, including prominent drugable signatures such as granulocyte and plasmablast/plasma cell activation. The substantial differences in gene expression between SLE and healthy individuals enables the classification of disease versus healthy status with median sensitivity and specificity of 83% and 100%, respectively. We explored the genetic regulation of blood transcriptome in SLE and found 3142 cis-expression quantitative trait loci (eQTLs). By integration of SLE genome-wide association study (GWAS) signals and eQTLs from 44 tissues from the Genotype-Tissue Expression (GTEx) consortium, we demonstrate that the genetic causality of SLE arises from multiple tissues with the top causal tissue being the liver, followed by brain basal ganglia, adrenal gland and whole blood. Collectively, our study defines distinct susceptibility and activity/severity signatures in SLE that may facilitate diagnosis, monitoring, and personalized therapy.

genomics

Deciphering cell lineage specification during male sex determination with single-cell RNA sequencing

The gonad is a unique biological system for studying cell fate decisions. However, major questions remain regarding the identity of somatic progenitor cells and the transcriptional events driving cell differentiation. Using time course single cell RNA sequencing on XY mouse gonads during sex determination, we identified a single population of somatic progenitor cells prior sex determination. A subset of these progenitors differentiate into Sertoli cells, a process characterized by a highly dynamic genetic program consisting of sequential waves of gene expression. Another subset of multipotent cells maintains their progenitor state but undergo significant transcriptional changes that restrict their competence towards a steroidogenic fate required for the differentiation of fetal Leydig cells. These results question the dogma of the existence of two distinct somatic cell lineages at the onset of sex determination and propose a new model of lineage specification from a unique progenitor cell population.

developmental biology

Hundreds of putative non-coding cis-regulatory drivers in chronic lymphocytic leukaemia and skin cancer

Perturbations of the coding genome and their role in cancer development have been studied extensively. However, the non-coding genomes contribution in cancer is poorly understood (1), not only because it is difficult to define the non-coding regulatory regions and the genes they regulate, but also because there is limited power owing to the regulatory regions small size. In this study, we try to resolve this issue by defining modules of coordinated non-coding regulatory regions of genes (Cis Regulatory Domains or CRDs). To do so, we use the correlation between histone modifications, assayed by ChIP-seq, in population samples of immortalized B-cells and skin fibroblasts. We screen for CRDs that accumulate an excess of somatic mutations in chronic lymphocytic leukaemia (CLL) and skin cancer, which affect these cell types, after accounting for somatic mutational patterns and biases. At 5% FDR, we find 90 CRDs with significant excess somatic of mutations in CLL, 60 of which regulate 126 genes, and in skin cancer 59 significant CRDs, 25 of which regulate 37 genes. The genes these CRDs regulate include ones already implicated in tumorigenesis, and are enriched in pathways already implicated in the respective cancers, like the B-cell receptor signalling pathway in CLL and the TGF{beta} signalling pathway in skin cancer. We discover that the somatic mutations in the significant CRDs of CLL are hitting bases more likely to be functional than the mutations in non-significant CRDs. Moreover, in both cancers, mutational signatures observed in the regulatory regions of significant CRDs deviate significantly from their null sequences. Both results indicate selection acting on CRDs during tumorigenesis. Finally, we find that the transcription factor biding sites that are disturbed by the somatic mutations in significant CRDs are enriched for factors known to be involved in cancer development. We are describing a new powerful approach to discover non-coding regions involved in tumorigenesis in CLL and skin cancer and this approach could be generalized to other cancers.

cancer biology