Search bioRxivSearch

Biology subjects

Porteous, D. J.

Publications and source records attributed to Porteous, D. J..

13 recordsLinked to original sources

Bayesian reassessment of the epigenetic architecture of complex traits

1Epigenetic DNA modification is partly under genetic control, and occurs in response to a wide range of environmental exposures. Linking epigenetic marks to clinical outcomes may provide greater insight into underlying molecular processes of disease, assist in the identification of therapeutic targets, and improve risk prediction. Here, we present a statistical approach, based on Bayesian inference, that estimates associations between disease risk and all measured epigenetic probes jointly, automatically controlling for both data structure (including cell-count effects, relatedness, and experimental batch effects) and correlations among probes. We benchmark our approach in simulation study, finding improved estimation of probe associations across a wide range of scenarios over existing approaches. Our method estimates the total proportion of disease risk captured by epigenetic probe variation, and when we applied it to measures of body mass index (BMI) and cigarette consumption behaviour in 5,101 individuals, we find that 66.7% (95% CI 60.0-72.8) of the variation in BMI and 67.7% (95% CI 58.4-76.9) of the variation in cigarette consumption can be captured by methylation array data from whole blood, independent of the variation explained by single nucleotide polymorphism markers. We find novel associations, with smoking behaviour associated with a methylation probe at the MNDA gene with >95% posterior inclusion probability, which is a myeloid cell nuclear differentiation antigen gene previously implicated as a biomarker for inflammation and non-Hodgkin lymphoma risk. We conduct unique genome-wide enrichment analyses, identifying blood cholesterol, lipid transport and sterol metabolism pathways for BMI, and response to xenobiotic stimulus and negative regulation of RNA polymerase II promoter transcription for smoking, all with >95% posterior inclusion probability of having methylation probes with associations >1.5 times larger than the average. Finally, we improve phenotypic prediction in two independent cohorts by 28.7% and 10.2% for BMI and smoking respectively over a LASSO model. These results imply that probe measures may capture large amounts of variance because they are likely a consequence of the phenotype rather than a cause. As a result, 'omics' data may enable accurate characterization of disease progression and identification of individuals who are on a path to disease. Our approach facilitates better understanding of the underlying epigenetic architecture of complex common disease and is applicable to any kind of genomics data.

genomics

Genome-wide meta-analysis of depression in 807,553 individuals identifies 102 independent variants with replication in a further 1,507,153 individuals

Major depression is a debilitating psychiatric illness that is typically associated with low mood, anhedonia and a range of comorbidities. Depression has a heritable component that has remained difficult to elucidate with current sample sizes due to the polygenic nature of the disorder. To maximise sample size, we meta-analysed data on 807,553 individuals (246,363 cases and 561,190 controls) from the three largest genome-wide association studies of depression. We identified 102 independent variants, 269 genes, and 15 gene-sets associated with depression, including both genes and gene-pathways associated with synaptic structure and neurotransmission. Further evidence of the importance of prefrontal brain regions in depression was provided by an enrichment analysis. In an independent replication sample of 1,306,354 individuals (414,055 cases and 892,299 controls), 87 of the 102 associated variants were significant following multiple testing correction. Based on the putative genes associated with depression this work also highlights several potential drug repositioning opportunities. These findings advance our understanding of the complex genetic architecture of depression and provide several future avenues for understanding aetiology and developing new treatment approaches.

genetics

The influence of X chromosome variants on trait neuroticism.

Autosomal variants have successfully been associated with trait neuroticism in genome-wide analysis of adequately-powered samples. But such studies have so far excluded the X chromosome from analysis. Here, we report genetic association analyses of X chromosome and XY pseudoautosomal single nucleotide polymorphisms (SNPs) and trait neuroticism using UK Biobank samples (N = 405,274). Significant association was found with neuroticism on the X chromosome for 204 markers found within three independent loci (a further 783 were suggestive). Most of these significant neuroticism-related X chromosome variants were located in intergenic regions (n = 713). Involvement of HS6ST2, which has been previously associated with sociability behaviour in the dog, was supported by single SNP and gene-based tests. We found that the amino acid and nucleotide sequences are highly conserved between dogs and humans. From the suggestive X chromosome variants, there were 19 nearby genes which could be linked to gene ontology information. Molecular function was primarily related to binding and catalytic activity; notable biological processes were cellular and metabolic, and nucleic acid binding and transcription factor protein classes were most commonly involved. X-variant heritability of neuroticism was estimated at 0.34% (SE = 0.07). A polygenic X-variant score created in an independent sample (maximum N {approx} 7300) did not predict significant variance in neuroticism, psychological distress, or depressive disorder. We conclude that the X chromosome harbours significant variants influencing neuroticism, and might prove important for other quantitative traits and complex disorders.

genetics

Epigenetic prediction of complex traits and death

BackgroundGenome-wide DNA methylation (DNAm) profiling has allowed for the development of molecular predictors for a multitude of traits and diseases. Such predictors may be more accurate than the self-reported phenotypes, and could have clinical applications. Here, penalised regression models were used to develop DNAm predictors for body mass index (BMI), smoking status, alcohol consumption, and educational attainment in a cohort of 5,100 individuals. Using an independent test cohort comprising 906 individuals, the proportion of phenotypic variance explained in each trait was examined for DNAm-based and genetic predictors. Receiver operator characteristic curves were generated to investigate the predictive performance of DNAm-based predictors, using dichotomised phenotypes. The relationship between DNAm scores and all-cause mortality (n = 214 events) was assessed via Cox proportional-hazards models.\n\nResultsThe DNAm-based predictors explained different proportions of the phenotypic variance for BMI (12%), smoking (60%), alcohol consumption (12%) and education (3%). The combined genetic and DNAm predictors explained 20% of the variance in BMI, 61% in smoking, 13% in alcohol consumption, and 6% in education. DNAm predictors for smoking, alcohol, and education but not BMI predicted mortality in univariate models. The predictors showed moderate discrimination of obesity (AUC=0.67) and alcohol consumption (AUC=0.75), and excellent discrimination of current smoking status (AUC=0.98). There was poorer discrimination of college-educated individuals (AUC=0.59).\n\nConclusionsDNAm predictors correlate with lifestyle factors that are associated with health and mortality. They may supplement DNAm-based predictors of age to identify the lifestyle profiles of individuals and predict disease risk.\n\nList of abbreviations

genomics

An epigenetic score for BMI based on DNA methylation correlates with poor physical health and major disease in the Lothian Birth Cohort 1936.

BackgroundThe relationship between obesity and adverse health is well established, but little is known about the contribution of DNA methylation to obesity-related health outcomes. Additionally, it is of interest whether such contributions are independent of those attributed by the most widely used clinical measure of body mass - the Body Mass Index (BMI).\n\nMethodWe tested whether an epigenetic BMI score accounts for inter-individual variation in health-related, cognitive, psychosocial and lifestyle outcomes in the Lothian Birth Cohort 1936 (n=903). Weights for the epigenetic BMI score were derived using penalised regression on methylation data from unrelated Generation Scotland participants (n=2566).\n\nResultsThe Epigenetic BMI score was associated with variables related to poor physical health (R2 ranges from 0.02-0.10), metabolic syndrome (R2 ranges from 0.01-0.09), lower crystallised intelligence (R2=0.01), lower health-related quality of life (R2=0.02), physical inactivity (R2=0.02), and social deprivation (R2=0.02). The epigenetic BMI score (per SD) was also associated with self-reported type 2 diabetes (OR 2.25, 95 % CI 1.74, 2.94), cardiovascular disease (OR 1.44, 95 % CI 1.23, 1.69) and high blood pressure (OR 1.21, 95% CI 1.13, 1.48; all at p<0.0011 after Bonferroni correction).\n\nConclusionsOur results show that regression models with epigenetic and phenotypic BMI scores as predictors account for a greater proportion of all outcome variables than either predictor alone, demonstrating independent and additive effects of epigenetic and phenotypic BMI scores.

epidemiology

DNA methylation age acceleration and risk factors for Alzheimer’s disease

INTRODUCTIONThe epigenetic clock is a DNA methylation-based estimate of biological age and is correlated with chronological age - the greatest risk factor for Alzheimers disease (AD). Genetic and environmental risk factors exist for AD, several of which are potentially modifiable. Here, we assess the relationship associations between the epigenetic clock and AD risk factors.\n\nMETHODSLinear mixed modelling was used to assess the relationship between age acceleration (the residual of biological age regressed onto chronological age) and AD risk factors relating to cognitive reserve, lifestyle, disease, and genetics in the Generation Scotland study (n=5,100).\n\nRESULTSWe report significant associations between the epigenetic clock and BMI, total:HDL cholesterol ratios, socioeconomic status, and smoking behaviour (Bonferroni-adjusted P<0.05).\n\nDISCUSSIONAssociations are present between environmental risk factors for AD and age acceleration. Measures to modify such risk factors might improve the risk profile for AD and the rate of biological ageing. Future longitudinal analyses are therefore warranted.

genomics

GWAS on family history of Alzheimer’s disease

Alzheimers disease (AD) is a public health priority for the 21st century. Risk reduction currently revolves around lifestyle changes with much research trying to elucidate the biological underpinnings. Using self-report of parental history of Alzheimers dementia for case ascertainment in a genome-wide association study of over 300,000 participants from UK Biobank (32,222 maternal cases, 16,613 paternal cases) and meta-analysing with published consortium data (n=74,046 with 25,580 cases across the discovery and replication analyses), six new AD-associated loci (P<5x10-8) are identified. Three contain genes relevant for AD and neurodegeneration: ADAM10, ADAMTS4, and ACE. Suggestive loci include drug targets such as VKORC1 (warfarin dose) and BZRAP1 (benzodiazepine receptor). We report evidence that association of SNPs and AD at the PVR gene is potentially mediated by both gene expression and DNA methylation in the prefrontal cortex. Our discovered loci may help to elucidate the biological mechanisms underlying AD and, given that many are existing drug targets for other diseases and disorders, warrant further exploration for potential precision medicine applications.

genetics

Inherited chromosomally integrated human herpesvirus 6 genomes are ancient, intact and potentially able to reactivate from telomeres

Human herpesviruses 6A and 6B (HHV6-A and HHV-6B; species Human herpesvirus 6A and Human herpesvirus 6B) have the capacity to integrate into telomeres, the essential capping structures of chromosomes that play roles in cancer and ageing. About 1% of people worldwide are carriers of chromosomally integrated HHV-6 (ciHHV-6), which is inherited as a genetic trait. Understanding the consequences of integration for the evolution of the viral genome, for the telomere and for the risk of disease associated with carrier status is hampered by a lack of knowledge about ciHHV-6 genomes. Here, we report an analysis of 28 ciHHV-6 genomes and show that they are significantly divergent from the few modern non-integrated HHV-6 strains for which complete sequences are currently available. In addition ciHHV-6B genomes in Europeans are more closely related to each other than to ciHHV-6B genomes from China and Pakistan, suggesting regional variation of the trait. Remarkably, at least one group of European ciHHV-6B carriers has inherited the same ciHHV-6B genome, integrated in the same telomere allele, from a common ancestor estimated to have existed 24,500 {+/-}10,600 years ago. Despite the antiquity of some, and possibly most, germline HHV-6 integrations, the majority of ciHHV-6B (95%) and ciHHV-6A (72%) genomes contain a full set of intact viral genes and therefore appear to have the capacity for viral gene expression and full reactivation.\n\nIMPORTANCEInheritance of HHV-6A or HHV-6B integrated into a telomere occurs at a low frequency in most populations studied to date but its characteristics are poorly understood. However, stratification of ciHHV-6 carriers in modern populations due to common ancestry is an important consideration for genome-wide association studies that aim to identify disease risks for these people. Here we present full sequence analysis of 28 ciHHV-6 genomes and show that ciHHV-6B in many carriers with European ancestry most likely originated from ancient integration events in a small number of ancestors. We propose that ancient ancestral origins for ciHHV-6A and ciHHV-6B are also likely in other populations. Moreover, despite their antiquity, all of the ciHHV-6 genomes appear to retain the capacity to express viral genes and most are predicted to be capable of full viral reactivation. These discoveries represent potentially important considerations in immune-compromised patients, in particular in organ transplantation and in stem cell therapy.

evolutionary biology

Meta-analysis of exome array data identifies six novel genetic loci for lung function

Over 90 regions of the genome have been associated with lung function to date, many of which have also been implicated in chronic obstructive pulmonary disease (COPD). We carried out meta-analyses of exome array data and three lung function measures: forced expiratory volume in one second (FEV1), forced vital capacity (FVC) and the ratio of FEV1 to FVC (FEV1/FVC). These analyses by the SpiroMeta and CHARGE consortia included 60,749 individuals of European ancestry from 23 studies, and 7,721 individuals of African Ancestry from 5 studies in the discovery stage, with follow-up in up to 111,556 independent individuals. We identified significant (P<2{middle dot}8x10-7) associations with six SNPs: a nonsynonymous variant in RPAP1, which is predicted to be damaging, three intronic SNPs (SEC24C, CASC17 and UQCC1) and two intergenic SNPs near to LY86 and FGF10. eQTL analyses found evidence for regulation of gene expression at three signals and implicated several genes including TYRO3 and PLAU. Further interrogation of these loci could provide greater understanding of the determinants of lung function and pulmonary disease.

genetics

Data Resource Profile: Generation Scotland Electronic Health Record

This paper provides the first detailed demonstration of the research value of the Electronic Health Record (EHR) linked to research data in Generation Scotland Scottish Family Health Study (GS:SFHS) participants, together with how to access this data. The structured, coded variables in the routine biochemistry, prescribing and morbidity records in particular represent highly valuable phenotypic data for a genomics research resource. Access to a wealth of other specialized datasets including cancer, mental health and maternity inpatient information is also possible through the same straightforward and transparent application process. The Electronic Health Record linked dataset is a key component of GS:SFHS, a biobank conceived in 1999 for the purpose of studying the genetics of health areas of current and projected public health importance. Over 24,000 adults were recruited from 2006 to 2011, with broad and enduring written informed consent for biomedical research. Consent was obtained from 23,603 participants for GS:SFHS study data to be linked to their Scottish National Health Service (NHS) records, using their Community Health Index (CHI) number. This identifying number is used for NHS Scotland procedures (registrations, attendances, samples, prescribing and investigations) and allows healthcare records for individuals to be linked across time and location. Here, we describe the NHS EHR dataset on the sub-cohort of 20,032 GS:SFHS participants with consent and mechanism for record linkage plus extensive genetic data. Together with existing study phenotypes, including family history and environmental exposures such as smoking, the EHR is a rich resource of real world data that can be used in research to characterise the health trajectory of participants, available at low cost and a high degree of timeliness, matched to DNA, urine and serum samples and genome-wide genetic information.

genetics

The Stratification Of Major Depressive Disorder Into Genetic Subgroups

Depression is a common and clinically heterogeneous mental health disorder that is frequently comorbid with other diseases and conditions. Stratification of depression may align sub-diagnoses more closely with their underling aetiology and provide more tractable targets for research and effective treatment. In the current study, we investigated whether genetic data could be used to identify subgroups within people with depression using the UK Biobank. Examination of cross-locus correlations was used to test for evidence of subgroups by examining whether there was clustering of independent genetic variants associated with eleven other complex traits and disorders in people with depression. We found evidence of a subgroup within depression using age of natural menopause variants (P = 1.69 x 10-3) and this effect remained significant in females (P = 1.18 x 10-3), but not males (P = 0.186). However, no evidence for this subgroup (P > 0.05) was found in Generation Scotland, iPSYCH, a UK Biobank replication cohort or the GERA cohort. In the UK Biobank, having depression was also associated with a later age of menopause (beta = 0.34, standard error = 0.06, P = 9.92 x 10-8). A potential age of natural menopause subgroup within depression and the association between depression and a later age of menopause suggests that they partially share a developmental pathway.

genetics

Genome-Wide Meta-Analyses Of Stratified Depression In Generation Scotland And UK Biobank

Few replicable genetic associations for Major Depressive Disorder (MDD) have been identified. However recent studies of depression have identified common risk variants by using either a broader phenotype definition in very large samples, or by reducing the phenotypic and ancestral heterogeneity of MDD cases. Here, a range of genetic analyses were applied to data from two large British cohorts, Generation Scotland and UK Biobank, to ascertain whether it is more informative to maximize the sample size by using data from all available cases and controls, or to use a refined subset of the data - stratifying by MDD recurrence or sex. Meta-analysis of GWAS data in males from these two studies yielded one genome-wide significant locus on 3p22.3. Three associated genes within this region (CRTAP, GLB1, and TMPPE) were significantly associated in subsequent gene-based tests. Meta-analyzed MDD, recurrent MDD and female MDD were each genetically correlated with 6 of 200 health-correlated traits, namely neuroticism, depressive symptoms, subjective well-being, MDD, a cross-disorder phenotype and Bipolar Disorder. Meta-analyzed male MDD showed no statistically significant correlations with these traits after correction for multiple testing. Whilst stratified GWAS analysis revealed a genome-wide significant locus for male MDD, the lack of independent replication, the equivalent SNP-based heritability estimates and the consistent pattern of genetic correlation with other health-related traits suggests that phenotypic stratification in currently available sample sizes is currently weakly justified. Based upon existing studies and our findings, the strategy of maximizing sample sizes is likely to provide the greater gain.

genetics

Genome-wide association study of alcohol consumption and genetic overlap with other health-related traits in UK Biobank (N=112,117).

Alcohol consumption has been linked to over 200 diseases and is responsible for over 5% of the global disease burden. Well known genetic variants in alcohol metabolizing genes, e.g. ALDH2, ADH1B, are strongly associated with alcohol consumption but have limited impact in European populations where they are found at low frequency. We performed a genome-wide association study (GWAS) of self-reported alcohol consumption in 112,117 individuals in the UK Biobank (UKB) sample of white British individuals. We report significant genome-wide associations at 8 independent loci. These include SNPs in alcohol metabolizing genes (ADH1B/ADH1C/ADH5) and 2 loci in KLB, a gene recently associated with alcohol consumption. We also identify SNPs at novel loci including GCKR, PXDN, CADM2 and TNFRSF11A. Gene-based analyses found significant associations with genes implicated in the neurobiology of substance use (CRHR1, DRD2), and genes previously associated with alcohol consumption (AUTS2). GCTA-GREML analyses found a significant SNP-based heritability of self-reported alcohol consumption of 13% (S.E.=0.01). Sex-specific analyses found largely overlapping GWAS loci and the genetic correlation between male and female alcohol consumption was 0.73 (S.E.=0.09, p-value = 1.37 x 10-16). Using LD score regression, genetic overlap was found between alcohol consumption and schizophrenia (rG=0.13, S.E=0.04), HDL cholesterol (rG=0.21, S.E=0.05), smoking (rG=0.49, S.E=0.06) and various anthropometric traits (e.g. Overweight, rG=-0.19, S.E.=0.05). This study replicates the association between alcohol consumption and alcohol metabolizing genes and KLB, and identifies 4 novel gene associations that should be the focus of future studies investigating the neurobiology of alcohol consumption.

genetics