Search bioRxivSearch

Biology subjects

Stegle, O.

Publications and source records attributed to Stegle, O..

At least 19 recordsLinked to original sources

Population-scale proteome variation in human induced pluripotent stem cells

Realising the potential of human induced pluripotent stem cell (iPSC) technology for drug discovery, disease modelling and cell therapy requires an understanding of variability across iPSC lines. While previous studies have characterized iPS cell lines genetically and transcriptionally, little is known about the variability of the iPSC proteome. Here, we present the first comprehensive proteomic iPSC dataset, analysing 202 iPSC lines derived from 151 donors. We characterise the major genetic determinants affecting proteome and transcriptome variation across iPSC lines and identify key regulatory mechanisms affecting variation in protein abundance. Our data identified >700 human iPSC protein quantitative trait loci (pQTLs). We mapped trans regulatory effects, identifying an important role for protein-protein interactions. We discovered that pQTLs show increased enrichment in disease-linked GWAS variants, compared with RNA-based eQTLs.

genomics

Cardelino: Integrating whole exomes and single-cell transcriptomes to reveal phenotypic impact of somatic variants

Decoding the clonal substructures of somatic tissues sheds light on cell growth, development and differentiation in health, ageing and disease. DNA-sequencing, either using bulk or using single-cell assays, has enabled the reconstruction of clonal trees from frequency and co-occurrence patterns of somatic variants. However, approaches to systematically characterize phenotypic and functional variations between individual clones are not established. Here we present cardelino (https://github.com/PMBio/cardelino), a computational method for inferring the clone of origin of individual cells that have been assayed using single-cell RNA-seq (scRNA-seq). After validating our model using simulations, we apply cardelino to matched scRNA-seq and exome sequencing data from 32 human dermal fibroblast lines, identifying hundreds of differentially expressed genes between cells from different somatic clones. These genes are frequently enriched for cell cycle and proliferation pathways, indicating a key role for cell division genes in non-neutral somatic evolution.\n\nKey findingsO_LIA novel approach for integrating DNA-seq and single-cell RNA-seq data to reconstruct clonal substructure for single-cell transcriptomes.\nC_LIO_LIEvidence for non-neutral evolution of clonal populations in human fibroblasts.\nC_LIO_LIProliferation and cell cycle pathways are commonly distorted in mutated clonal populations.\nC_LI

genomics

Kipoi: accelerating the community exchange and reuse of predictive models for genomics

Advanced machine learning models applied to large-scale genomics datasets hold the promise to be major drivers for genome science. Once trained, such models can serve as a tool to probe the relationships between data modalities, including the effect of genetic variants on phenotype. However, lack of standardization and limited accessibility of trained models have hampered their impact in practice. To address this, we present Kipoi, a collaborative initiative to define standards and to foster reuse of trained models in genomics. Already, the Kipoi repository contains over 2,000 trained models that cover canonical prediction tasks in transcriptional and post-transcriptional gene regulation. The Kipoi model standard grants automated software installation and provides unified interfaces to apply and interpret models. We illustrate Kipoi through canonical use cases, including model benchmarking, transfer learning, variant effect prediction, and building new models from existing ones. By providing a unified framework to archive, share, access, use, and build on models developed by the community, Kipoi will foster the dissemination and use of machine learning models in genomics.

bioinformatics

Genome-scale oscillations in DNA methylation during exit from pluripotency

Pluripotency is accompanied by the erasure of parental epigenetic memory with naive pluripotent cells exhibiting global DNA hypomethylation both in vitro and in vivo. Exit from pluripotency and priming for differentiation into somatic lineages is associated with genome-wide de novo DNA methylation. We show that during this phase, coexpression of enzymes required for DNA methylation turnover, DNMT3s and TETs, promotes cell-to-cell variability in this epigenetic mark. Using a combination of single-cell sequencing and quantitative biophysical modelling, we show that this variability is associated with coherent, genome-scale, oscillations in DNA methylation with an amplitude dependent on CpG density. Analysis of parallel single-cell transcriptional and epigenetic profiling provides evidence for oscillatory dynamics both in vitro and in vivo. These observations provide fresh insights into the emergence of epigenetic heterogeneity during early embryo development, indicating that dynamic changes in DNA methylation might influence early cell fate decisions.\n\nHighlightsO_LICo-expression of DNMT3s and TETs drive genome-scale oscillations of DNA methylation\nC_LIO_LIOscillation amplitude is greatest at a CpG density characteristic of enhancers\nC_LIO_LICell synchronisation reveals oscillation period and link with primary transcripts\nC_LIO_LIMultiomic single-cell profiling provides evidence for oscillatory dynamics in vivo\nC_LI

genomics

Tandem duplications lead to loss of fitness effects in CRISPR-Cas9 data

CRISPR-Cas9 gene-editing is widely used to study gene function and is being advanced for therapeutic applications. Structural rearrangements are a ubiquitous feature of cancers and their impact on CRISPR-Cas9 gene-editing has not yet been systematically assessed. Utilising CRISPR-Cas9 knockout screens for 163 cancer cell lines, we demonstrate that targeting tandem amplified regions is highly detrimental to cellular fitness, in contrast to amplifications caused by chromosomal duplications which have little to no effect. Genomically clustered Cas9 double-strand DNA breaks are associated with a strong gene-independent decrease in cell fitness. We systematically identified collateral vulnerabilities in 25% of cancer cells, introduced by tandem amplifications of tissue non-expressed genes. Our analysis demonstrates the importance of structural rearrangements in mediating the effect of CRISPR-Cas9-induced DNA damage, with implications for the use of CRISPR-Cas9 gene-editing technology, and how resulting collateral vulnerabilities are a generalisable strategy to target cancer cells.

genomics

Combined single cell profiling of expression and DNA methylation reveals splicing regulation and heterogeneity

BackgroundAlternative splicing is a key regulatory mechanism in eukaryotic cells and increases the effective number of functionally distinct gene products. Using bulk RNA sequencing, splicing variation has been studied across human tissues and in genetically diverse populations. This has identified disease-relevant splicing events, as well as associations between splicing and genomic variations, including sequence composition and conservation. However, variability in splicing between single cells from the same tissue or cell type and its determinants remain poorly understood.\n\nResultsWe applied parallel DNA methylation and transcriptome sequencing to differentiating human induced pluripotent stem cells to characterize splicing variation (exon skipping) and its determinants. Our results shows that variation in single-cell splicing can be accurately predicted based on local sequence composition and genomic features. We observe moderate but consistent contributions from local DNA methylation profiles to splicing variation across cells. A combined model that is built based on sequence as well as DNA methylation information accurately predicts different splicing modes of individual cassette exons (AUC=0.85). These categories include the conventional inclusion and exclusion patterns, but also more subtle modes of cell-to-cell variation in splicing. Finally, we identified and characterized associations between DNA methylation and splicing changes during cell differentiation.\n\nConclusionsOur study yields new insights into alternative splicing at the single-cell level and reveals a previously underappreciated link between DNA methylation variation and splicing.

bioinformatics

Genome-wide association study of social genetic effects on 170 phenotypes in laboratory mice

The phenotype of one individual can be affected not only by the individuals own genotypes (direct genetic effects, DGE) but also by genotypes of interacting partners (indirect genetic effects, IGE). IGE have been detected using polygenic models in multiple species, including laboratory mice and humans. However, the underlying mechanisms remain largely unknown. Genome-wide association studies of IGE (igeGWAS) can point to IGE genes, but have not yet been applied to non-familial IGE arising from "peers" and affecting biomedical phenotypes. In addition, the extent to which igeGWAS will identify loci not identified by dgeGWAS remains an open question. Finally, findings from igeGWAS have not been confirmed by experimental manipulation. We leveraged a dataset of 170 behavioural, physiological and morphological phenotypes measured in 1,812 genetically heterogeneous laboratory mice to study IGE arising between same-sex, adult, unrelated laboratory mice housed in the same cage. We developed methods for igeGWAS in this context and identified 24 significant IGE loci for 17 phenotypes (FDR < 10%). There was no overlap between IGE loci and DGE loci for the same phenotype, which was consistent with the moderate genetic correlations between DGE and IGE for the same phenotype estimated using polygenic models. Finally, we fine-mapped seven significant IGE loci to individual genes and confirmed, in an experiment with a knockout model, that Epha4 gives rise to IGE on stress-coping strategy and wound healing. Our results demonstrate the potential for igeGWAS to identify IGE genes and shed some light into the mechanisms of peer influence.

genetics

Identifying the genetic basis of variation in cell behaviour in human iPS cell lines from healthy donors

Large cohorts of human iPSCs from healthy donors are potentially a powerful tool for investigating the relationship between genetic variants and cellular phenotypes. Here we integrate high content imaging, gene expression and DNA sequence datasets for over 100 human iPSC lines to identify the genetic basis of inter-individual variability in cell behaviour. By applying a dimensionality reduction approach, Probabilistic Estimation of Expression Residuals (PEER), we identified genes that correlated in expression with intrinsic (genetic) and extrinsic (ECM) factors. However, variation in mRNA levels could not account for outlier cell behaviour. Instead, we identified rare, deleterious SNVs in the coding sequence of genes involved in ECM adhesion that occurred in cell lines that were outliers for one or more phenotypes such as cell spreading. These also correlated with altered germ layer differentiation on micropatterned surfaces. Our study thus establishes a strategy for integrating genetic and cell biological measurements for high-throughput analysis.

cell biology

A linear mixed model approach to study multivariate gene-environment interactions

Different environmental factors, including diet, physical activity, or external conditions can contribute to genotype-environment interactions (GxE). Although high-dimensional environmental data are increasingly available, and multiple environments have been implicated with GxE at the same loci, multi-environment tests for GxE are not established. Such joint analyses can increase power to detect GxE and improve the interpretation of these effects. Here, we propose the structured linear mixed model (StructLMM), a computationally efficient method to test for and characterize loci that interact with multiple environments. After validating our model using simulations, we apply StructLMM to body mass index in UK Biobank, where our method detects previously known and novel GxE signals. Finally, in an application to a large blood eQTL dataset, we demonstrate that StructLMM can be used to study interactions with hundreds of environmental variables.

genetics

Modelling cell-cell interactions from spatial molecular data with spatial variance component analysis

Technological advances allow for assaying multiplexed spatially resolved RNA and protein expression profiling of individual cells, thereby capturing physiological tissue contexts of single cell variation. While methods for the high-throughput generation of spatial expression profiles are increasingly accessible, computational methods for studying the relevance of the spatial organization of tissues on cell-cell heterogeneity are only beginning to emerge. Here, we present spatial variance component analysis (SVCA), a computational framework for the analysis of spatial molecular data. SVCA enables quantifying the effect of cell-cell interactions, as well as environmental and intrinsic cell features on the expression levels of individual genes or proteins. In application to a breast cancer Imaging Mass Cytometry dataset, our model allows for robustly estimating spatial variance signatures, identifying cell-cell interactions as a major driver of expression heterogeneity. Finally, we apply SVCA to high-dimensional imaging-derived RNA data, where we identify molecular pathways that are linked to cell-cell interactions.

bioinformatics

LiMMBo: a simple, scalable approach for linear mixed models in high-dimensional genetic association studies

Genome-wide association studies have helped to shed light on the genetic architecture of complex traits and diseases. Deep phenotyping of population cohorts is increasingly applied, where multi-to high-dimensional phenotypes are recorded in the individuals. Whilst these rich datasets provide important opportunities to analyse complex trait structures and pleiotropic effects at a genome-wide scale, existing statistical methods for joint genetic analyses are hampered by computational limitations posed by high-dimensional phenotypes. Consequently, such multivariate analyses are currently limited to a moderate number of traits. Here, we introduce a method that combines linear mixed models with bootstrapping (LiMMBo) to enable computationally efficient joint genetic analysis of high-dimensional phenotypes. Our method builds on linear mixed models, thereby providing robust control for population structure and other confounding factors, and the model scales to larger datasets with up to hundreds of phenotypes. We first validate LiMMBo using simulations, demonstrating consistent covariance estimates at greatly reduced computational cost compared to existing methods. We also find LiMMBo yields consistent power advantages compared to univariate modelling strategies, where the advantages of multivariate mapping increases substantially with the phenotype dimensionality. Finally, we applied LiMMBo to 41 yeast growth traits to map their genetic determinants, finding previously known and novel pleiotropic relationships in this high-dimensional phenotype space. LiMMBo is accessible as open source software (https://github.com/HannahVMeyer/limmbo).\n\nAuthor summaryIn multi-trait genetic association studies one is interested in detecting genetic variants that are associated with one or multiple traits. Genetic variants that influence two or more traits are referred to as pleiotropic. Multivariate linear mixed models have been successfully applied to detect pleiotropic effects, by jointly modelling association signals across traits. However, these models are currently limited to a moderate number of phenotypes as the number of model parameters grows steeply with the number of phenotypes, raising a computational burden. We developed LiMMBo, a new approach for the joint analysis of high-dimensional phenotypes. Our method reduces the number of effective model parameters by introducing an intermediate subsampling step. We validate this strategy using simulations, where we apply LiMMBo for the genetic analysis of hundreds of phenotypes, detecting pleiotropic effects for a wide range of simulated genetic architectures. Finally, to illustrate LiMMBo in practice, we apply the model to a study of growth traits in yeast, where we identify pleiotropic effects for traits with formerly known genetic effects as well as revealing previously unconnected traits.

genomics

Assessing the Gene Regulatory Landscape in 1,188 Human Tumors

Cancer is characterised by somatic genetic variation, but the effect of the majority of non-coding somatic variants and the interface with the germline genome are still unknown. We analysed the whole genome and RNA-Seq data from 1,188 human cancer patients as provided by the Pan-cancer Analysis of Whole Genomes (PCAWG) project to map cis expression quantitative trait loci of somatic and germline variation and to uncover the causes of allele-specific expression patterns in human cancers. The availability of the first large-scale dataset with both whole genome and gene expression data enabled us to uncover the effects of the non-coding variation on cancer. In addition to confirming known regulatory effects, we identified novel associations between somatic variation and expression dysregulation, in particular in distal regulatory elements. Finally, we uncovered links between somatic mutational signatures and gene expression changes, including TERT and LMO2, and we explained the inherited risk factors in APOBEC-related mutational processes. This work represents the first large-scale assessment of the effects of both germline and somatic genetic variation on gene expression in cancer and creates a valuable resource cataloguing these effects.

cancer biology

Multi-Omics factor analysis disentangles heterogeneity in blood cancer

Multi-omic studies promise the improved characterization of biological processes across molecular layers. However, methods for the unsupervised integration of the resulting heterogeneous datasets are lacking. We present Multi-Omics Factor Analysis (MOFA), a computational method for discovering the principal sources of variation in multi-omic datasets. MOFA infers a set of (hidden) factors that capture biological and technical sources of variability. It disentangles axes of heterogeneity that are shared across multiple modalities and those specific to individual data modalities. The learnt factors enable a variety of downstream analyses, including identification of sample subgroups, data imputation, and the detection of outlier samples. We applied MOFA to a cohort of 200 patient samples of chronic lymphocytic leukaemia, profiled for somatic mutations, RNA expression, DNA methylation and ex-vivo drug responses. MOFA identified major dimensions of disease heterogeneity, including immunoglobulin heavy chain variable region status, trisomy of chromosome 12 and previously underappreciated drivers, such as response to oxidative stress. In a second application, we used MOFA to analyse single-cell multiomics data, identifying coordinated transcriptional and epigenetic changes along cell differentiation.

bioinformatics

A pan cancer analysis of promoter activity highlights the regulatory role of alternative transcription start sites and their association with noncoding mutations

Most human protein-coding genes are regulated by multiple, distinct promoters, suggesting that the choice of promoter is as important as its level of transcriptional activity. While the role of promoters as driver elements in cancer has been recognized, the contribution of alternative promoters to regulation of the cancer transcriptome remains largely unexplored. Here we infer active promoters using RNA-Seq data from 1,188 cancer samples with matched whole genome sequencing data. We find that alternative promoters are a major contributor to context-specific regulation of isoform expression and that alternative promoters are frequently deregulated in cancer, affecting known cancer-genes and novel candidates. Our study suggests that a highly dynamic landscape of active promoters shapes the cancer transcriptome, opening many opportunities to further explore the interplay of regulatory mechanism and noncoding somatic mutations with transcriptional aberrations in cancer.

genomics

SpatialDE - Identification Of Spatially Variable Genes

Technological advances have enabled low-input RNA-sequencing, paving the way for assaying transcriptome variation in spatial contexts, including in tissues. While the generation of spatially resolved transcriptome maps is increasingly feasible, computational methods for analysing the resulting data are not established. Existing analysis strategies either ignore the spatial component of gene expression variation, or require discretization of the cells into coarse grained groups.\n\nTo address this, we have developed SpatialDE, a computational framework for identifying and characterizing spatially variable genes. Our method generalizes variable gene selection, as used in population-and single-cell studies, to spatial expression profiles. To illustrate the broad utility of our approach, we apply SpatialDE to spatial transcriptomics data, and to data from single cell methods based on multiplexed in situ hybridisation (SeqFISH and MERFISH). SpatialDE enables the statistically robust identification of spatially variable genes, thereby identifying genes with known disease implications, several of which are missed by conventional variable gene selection. Additionally, to enable gene-expressed based histology, SpatialDE implements a spatial gene clustering model which we call \"automatic expression histology,\" allowing to classify genes into groups with distinct spatial patterns.

genomics

Joint Profiling Of Chromatin Accessibility, DNA Methylation And Transcription In Single Cells

Parallel single-cell sequencing protocols represent powerful methods for investigating regulatory relationships, including epigenome-transcriptome interactions. Here, we report a novel single-cell method for parallel chromatin accessibility, DNA methylation and transcriptome profiling. scNMT-seq (single-cell nucleosome, methylation and transcription sequencing) uses a GpC methyltransferase to label open chromatin followed by bisulfite and RNA sequencing. We validate scNMT-seq by applying it to differentiating mouse embryonic stem cells, finding links between all three molecular layers and revealing dynamic coupling between epigenomic layers during differentiation.

genomics

Multi-Tissue DNA Methylation Age Predictor In Mouse

BackgroundDNA-methylation changes at a discrete set of sites in the human genome are predictive of chronological and biological age. However, it is not known whether these changes are causative or a consequence of an underlying ageing process. It has also not been shown whether this epigenetic clock is unique to humans or conserved in the more experimentally tractable mouse.\n\nResultsWe have generated a comprehensive set of genome-scale base-resolution methylation maps from multiple mouse tissues spanning a wide range of ages. Many CpG sites show significant tissue-independent correlations with age and allowed us to develop a multi-tissue predictor of age in the mouse. Our model, which estimates age based on DNA methylation at 329 unique CpG sites, has a median absolute error of 3.33 weeks, and has similar properties to the recently described human epigenetic clock. Using publicly available datasets, we find that the mouse clock is accurate enough to measure effects on biological age, including in the context of interventions. While females and males show no significant differences in predicted DNA methylation age, ovariectomy results in significant age acceleration in females. Furthermore, we identify significant differences in age-acceleration dependent on the lipid content of the offspring diet.\n\nConclusionsHere we identify and characterize an epigenetic predictor of age in mice, the mouse epigenetic clock. This clock will be instrumental for understanding the biology of ageing and will allow modulation of its ticking rate and resetting the clock in vivo to study the impact on biological age.

genomics

Interactions between genetic variation and cellular environment in skeletal muscle gene expression

From whole organisms to individual cells, responses to environmental conditions are influenced by genetic makeup, where the effect of genetic variation on a trait depends on the environmental context. RNA-sequencing quantifies gene expression as a molecular trait, and is capable of capturing both genetic and environmental effects. In this study, we explore opportunities of using allele-specific expression (ASE) to discover cis acting genotype-environment interactions (GxE) - genetic effects on gene expression that depend on an environmental condition. Treating 17 common, clinical traits as approximations of the cellular environment of 267 skeletal muscle biopsies, we identify 10 candidate interaction quantitative trait loci (iQTLs) across 6 traits (12 unique gene-environment trait pairs; 10% FDR per trait) including sex, systolic blood pressure, and low-density lipoprotein cholesterol. Although using ASE is in principle a promising approach to detect GxE effects, replication of such signals can be challenging as validation requires harmonization of environmental traits across cohorts and a sufficient sampling of heterozygotes for a transcribed SNP. Comprehensive discovery and replication will require large human transcriptome datasets, or the integration of multiple transcribed SNPs, coupled with standardized clinical phenotyping.

genetics