Search bioRxivSearch

Biology subjects

Teschendorff, A.

Publications and source records attributed to Teschendorff, A..

6 recordsLinked to original sources

ebayGSEA: An improved Gene Set Enrichment Analysis method for Epigenome-Wide-Association Studies

MotivationThe biological interpretation of differentially methylated sites derived from Epigenome-Wide-Association Studies remains a significant challenge. Gene Set Enrichment Analysis (GSEA) is a general tool to help aid biological interpretation, yet its correct and unbiased implementation in the EWAS context is difficult due to the differential probe representation of Illumina Infinium DNA methylation beadchips.\n\nResultsWe present a novel GSEA method, called ebayGSEA, which ranks genes, not CpGs, according to the overall level of differential methylation, as assessed using all the probes mapping to the given gene. Applied on simulated and real EWAS data, we show how ebayGSEA may exhibit higher sensitivity and specificity than the current state-of-the-art, whilst also avoiding differential probe representation bias. Thus, ebayGSEA will be a useful additional tool to aid the interpretation of EWAS data.\n\nAvailability and implementationebayGSEA is available from https://github.com/aet21/ebayGSEA, and has been incorporated into the ChAMP Bioconductor package (https://www.bioconductor.org).

bioinformatics

Identification of differentially methylated cell-types in Epigenome-Wide Association Studies

An outstanding challenge of Epigenome-Wide Association Studies (EWAS) performed in complex tissues is the identification of the specific cell-type(s) responsible for the observed differential DNA methylation. Here, we present a novel statistical algorithm, called CellDMC, which is able to identify not only differentially methylated positions, but also the specific cell-type(s) driving the differential methylation. We provide extensive validation of CellDMC on in-silico mixtures of DNA methylation data generated with different technologies, as well as on real mixtures from epigenome-wide-association and cancer epigenome studies. We demonstrate how CellDMC can achieve over 90% sensitivity and specificity in scenarios where current state-of-the-art methods fail to identify differential methylation. By applying CellDMC to a smoking EWAS performed in buccal swabs, we identify differentially methylated positions occurring in the epithelial compartment, which we validate in smoking-related lung cancer. CellDMC may help towards the identification of causal DNA methylation alterations in disease.

bioinformatics

Tensorial Blind Source Separation for Improved Analysis of Multi-Omic Data

There is an increased need for integrative analyses of multi-omic data. Although several algorithms for analysing multi-omic data exist, no study has yet performed a detailed comparison of these methods in biologically relevant contexts. Here we benchmark a novel tensorial independent component analysis (tICA) algorithm against current state-of-the-art methods. Using simulated and real multi-omic data, we find that tICA outperforms established methods in identifying biological sources of data variation at a significantly reduced computational cost. Using two independent multi cell-type EWAS, we further demonstrate how tICA can identify, in the absence of genotype information, mQTLs at a higher sensitivity than competing multi-way algorithms. We validate mQTLs found with tICA in an independent set, and demonstrate that approximately 75% of mQTLs are independent of blood cell subtype. In an application to multi-omic cancer data, tICA identifies many gene modules whose expression variation across tumors is driven by copy number or DNA methylation changes, but whose deregulation relative to the normal state is independent such alterations, an important finding that we confirm by direct analysis of individual data types. In summary, tICA is a powerful novel algorithm for decomposing multi-omic data, which will be of great value to the research community.

bioinformatics

Appraising the causal relevance of DNA methylation for risk of lung cancer

DNA methylation changes in peripheral blood have been identified in relation to lung cancer risk. However, the causal nature of these associations remains to be fully elucidated. Meta-analysis of four epigenome-wide association studies (918 cases, 918 controls) revealed differential methylation at 16 CpG sites (FDR < 0.05) in relation to lung cancer risk. A two-sample Mendelian randomization analysis, using genetic instruments for methylation at 14 of the 16 CpG sites, and 29,863 cases and 55,586 controls from the TRICL-ILCCO lung cancer consortium, was performed to appraise the causal role of methylation at these sites on lung cancer. This approach provided little evidence that DNA methylation in peripheral blood at the 14 CpG sites play a causal role in lung cancer development, including for cg05575921 AHRR, where methylation is strongly associated with lung cancer risk. Further studies are needed to investigate the causal role played by DNA methylation in lung tissue.

epidemiology

Quantifying Waddington’s epigenetic landscape: a comparison of single-cell potency measures

Over 60 years ago Waddington proposed an epigenetic landscape model of cellular differentiation, whereby cell-fate transitions are modelled as canalization events, with stable cell states occupying the basins or attractor states1, 2. A key ingredient of this landscape is the energy potential, or height3, which correlates with cell-potency. To date, very few explicit biophysical models for estimating single-cell potency have been proposed. Using 9 independent experiments, encompassing over 6,600 high-quality single-cell RNA-Seq profiles, we here demonstrate that single-cell potency can be approximated as the graph entropy of a Markov Chain process on a model signaling network. Our analysis highlights that other proposed single-cell potency measures are not robust, whilst also revealing that integration with orthogonal systems-level information improves potency estimates. Thus, this study provides a foundation for an improved systems-level understanding of single-cell potency, which may have profound implications for the discovery of novel stem-and progenitor cell phenotypes.

bioinformatics

Correcting For Cell-Type Heterogeneity In Epigenome-Wide Association Studies: Premature Analyses And Conclusions

Recently, a study by Rahmani et al [1] claimed that a reference-free cell-type deconvolution method, called ReFACTor, leads to improved power and improved estimates of cell-type composition compared to competing reference-free and reference-based methods in the context of Epigenome-Wide Association Studies (EWAS). However, we identified many critical flaws (both conceptual and statistical in nature), which seriously question the validity of their claims. We outlined constructive criticism in a recent correspondence letter, Zheng et al [2]. The purpose of this letter is two-fold. First, to present additional analyses, which demonstrate that our original criticism is statistically sound. Second, to highlight additional serious concerns, which Rahmani et al have not yet addressed. In summary, we find that ReFACTor has not been demonstrated to outperform state-of-the-art reference-free methods such as SVA or RefFreeEWAS, nor state-of-the-art reference-based methods. Thus, the claim by Rahmani et al (a claim reiterated in their recent response letter [3]) that ReFACT or represents an advance over the state-of-the-art is not supported by an objective and rigorous statistical analysis of the data.

bioinformatics