Search bioRxivSearch

Biology subjects

Dahir, A.

Publications and source records attributed to Dahir, A..

2 recordsLinked to original sources

Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation

BackgroundThe combination of sodium bisulfite treatment with highly-parallel sequencing is a common method for quantifying DNA methylation across the genome. The power to detect between-group differences in DNA methylation using bisulfite-sequencing approaches is influenced by both experimental (e.g. read depth, missing data and sample size) and biological (e.g. mean level of DNA methylation and difference between groups) parameters. There is, however, no consensus about the optimal thresholds for filtering bisulfite sequencing data with implications for the reproducibility of findings in epigenetic epidemiology. ResultsWe used a large reduced representation bisulfite sequencing (RRBS) dataset to assess the distribution of read depth across DNA methylation sites and the extent of missing data. To investigate how various study variables influence power to identify DNA methylation differences between groups, we developed a framework for simulating bisulfite sequencing data. As expected, sequencing read depth, group size, and the magnitude of DNA methylation difference between groups all impacted upon statistical power. The influence on power was not dependent on one specific parameter, but reflected the combination of study-specific variables. As a resource to the community, we have developed a tool, POWEREDBiSeq, which utilizes our simulation framework to predict study-specific power for the identification of DNAm differences between groups, taking into account user-defined read depth filtering parameters and the minimum sample size per group. ConclusionsOur data-driven approach highlights the importance of filtering bisulfite-sequencing data by minimum read depth and illustrates how the choice of threshold is influenced by the specific study design and the expected differences between groups being compared. The POWEREDBiSeq tool can help users identify the level of data filtering needed to optimize power and aims to improve the reproducibility of bisulfite sequencing studies.

bioinformatics

Recalibrating the Epigenetic Clock: Implications for Assessing Biological Age in the Human Cortex

Human DNA-methylation data have been used to develop biomarkers of ageing - referred to epigenetic clocks - that have been widely used to identify differences between chronological age and biological age in health and disease including neurodegeneration, dementia and other brain phenotypes. Existing DNA methylation clocks are highly accurate in blood but are less precise when used in older samples or on brain tissue. We aimed to develop a novel epigenetic clock that performs optimally in human cortex tissue and has the potential to identify phenotypes associated with biological ageing in the brain. We generated an extensive dataset of human cortex DNA methylation data spanning the life-course (n = 1,397, ages = 1 to 104 years). This dataset was split into training and testing samples (training: n = 1,047; testing: n = 350). DNA methylation age estimators were derived using a transformed version of chronological age on DNA methylation at specific sites using elastic net regression, a supervised machine learning method. The cortical clock was subsequently validated in a novel human cortex dataset (n = 1,221, ages = 41 to 104 years) and tested for specificity in a large whole blood dataset (n = 1,175, ages = 28 to 98 years). We identified a set of 347 DNA methylation sites that, in combination optimally predict age in the human cortex. The sum of DNA methylation levels at these sites weighted by their regression coefficients provide the cortical DNA methylation clock age estimate. The novel clock dramatically out-performed previously reported clocks in additional cortical datasets. Our findings suggest that previous associations between predicted DNA methylation age and neurodegenerative phenotypes might represent false positives resulting from clocks not robustly calibrated to the tissue being tested and for phenotypes that become manifest in older ages. The age distribution and tissue type of samples included in training datasets need to be considered when building and applying epigenetic clock algorithms to human epidemiological or disease cohorts.

genomics