Search bioRxivSearch

Biology subjects

Campbell, A.

Publications and source records attributed to Campbell, A..

12 recordsLinked to original sources

VEP-G2P: A Tool for Efficient, Flexible and Scalable Diagnostic Filtering of Genomic Variants

PurposeWe aimed to develop an efficient, flexible, scalable and evidence-based approach to sequence-based diagnostic analysis/re-analysis of conditions with very large numbers of different causative genes. We then wished to define the expected rate of plausibly causative variants coming through strict filtering in control in comparison to disease populations to quantify background diagnostic \"noise\".\n\nMethodsWe developed G2P (www.ebi.ac.uk/gene2phenotype) as an online system to facilitate the development, validation, curation and distribution of large-scale, evidence-based datasets for use in diagnostic variant filtering. Each locus-genotype-mechanism-disease-evidence thread (LGMDET) associates an allelic requirement and a mutational consequence at a defined locus with a disease entity and a confidence level and evidence links. We then developed an extension to Ensembl Variant Effect Predictor (VEP), VEP-G2P, which can filter based on G2P other widely used gene panel curation systems. We compared the output of disease-associated and control whole exome sequence (WES) using Developmental Disorders G2P (G2PDD; 2044 LGMDETs) and constitutional cancer predisposition G2P (G2PCancer; 128 LGMDETs).\n\nResultsWe have shown a sensitivity/precision of 97.3%/33% and 81.6%/22.7% for causative de novo and inherited variants respectively using VEP-G2PDD in DDD study probands WES. Many of the apparently diagnostic genotypes \"missed\" are likely false-positive reports with lower minor allele frequencies and more severe predicted consequences being diagnostically-discriminative features.\n\nConclusionCase:control comparisons using VEP-G2PDD established an observed:expected ratio of 1:30,000 plausibly causative variants in proband WES to ~1:40,000 reportable but presumed-benign variants in controls. At least half the filtered variants in probands represent background \"noise\". Supporting phenotypic evidence is, therefore, necessary in genetically-heterogeneous disorders. G2P and VEP-G2P provides a practical approach to optimize disease-specific filtering parameters in diagnostic genetic research.

genetics

Large-scale genome-wide association meta-analysis of endometriosis reveals 13 novel loci and genetically-associated comorbidity with other pain conditions

Endometriosis is a common complex inflammatory condition characterised by the presence of endometrium-like tissue outside the uterus, mainly in the pelvic area. It is associated with chronic pelvic pain and infertility, and its pathogenesis remains poorly understood. The disease is typically classified according to the revised American Fertility Society (rAFS) 4-stage surgical assessment system, although stage does not correlate well with symptomatology or prognosis. Previously identified genetic variants mainly are associated with stage III/IV disease, highlighting the need for further phenotype-stratified analysis that requires larger datasets. We conducted a meta-analysis of 15 genome-wide association studies (GWAS) and a replication analysis, including 58,115 cases and 733,480 controls in total, and sub-phenotype analyses of stage I/II, stage III/IV and infertility-associated endometriosis cases. This revealed 27 genetic loci associated with endometriosis at the genome-wide p-value threshold (P<5x10-8), 13 of which are novel and an additional 8 novel genes identified from gene-based association analyses. Of the 27 loci, 21 (78%) had greater effect sizes in stage III/IV disease compared to stage I/II, 1 (4%) had greater effect size in stage I/II compared to stage III/IV and 17 (63%) had greater effect sizes when restricted to infertility-associated endometriosis cases compared to overall endometriosis. These results suggest that specific variants may confer risk for different sub-types of endometriosis through distinct pathways. Analyses of genetic variants underlying different pain symptoms reported in the UK Biobank showed that 7/9 had positive significant (p<1.28x103) positive genetic correlations with endometriosis, suggesting a genetic basis for sensitivity to pain in general. Additional conditions with significant positive genetic correlations with endometriosis included uterine fibroids, excessive and irregular menstrual bleeding, osteoarthritis, diabetes as well as menstrual cycle length and age at menarche. These results provide a basis for fine-mapping of the causal variants at these 27 loci, and for functional follow-up to understand their contribution to endometriosis and its potential subtypes.

genomics

Epigenetic signatures of starting and stopping smoking

BackgroundMultiple studies have made robust associations between differential DNA methylation and exposure to cigarette smoke. But whether a DNA methylation phenotype is established immediately upon exposure, or only after prolonged exposure is less well-established. Here, we assess DNA methylation patterns in current smokers in response to dose and duration of exposure, along with the effects of smoking cessation on DNA methylation in former smokers.\n\nMethodsDimensionality reduction was applied to DNA methylation data at 90 previously identified smoking-associated CpG sites for over 4,900 individuals in the Generation Scotland cohort. K-means clustering was performed to identify clusters associated with current and never smoker status based on these methylation patterns. Cluster assignments were assessed with respect to duration of exposure in current smokers (years as a smoker), time since smoking cessation in former smokers (years), and dose (cigarettes per day).\n\nResultsTwo clusters were specified, corresponding to never smokers (97.5% of whom were assigned to Cluster 1) and current smokers (81.1% of whom were assigned to Cluster 2). The exposure time point from which >50% of current smokers were assigned to the smoker-enriched cluster varied between 5-9 years in heavier smokers and between 15-19 years in lighter smokers. Low-dose former smokers were more likely to be assigned to the never smoker-enriched cluster from the first year following cessation. In contrast, a period of at least two years was required before the majority of former high-dose smokers were assigned to the never smoker-enriched cluster.\n\nConclusionsOur findings suggest that smoking-associated DNA methylation changes are a result of prolonged exposure to cigarette smoke, and can be reversed following cessation. The length of time in which these signatures are established and recovered is dose dependent. Should DNA methylation-based signatures of smoking status be predictive of smoking-related health outcomes, our findings may provide an additional criterion on which to stratify risk.

genomics

Epigenetic prediction of complex traits and death

BackgroundGenome-wide DNA methylation (DNAm) profiling has allowed for the development of molecular predictors for a multitude of traits and diseases. Such predictors may be more accurate than the self-reported phenotypes, and could have clinical applications. Here, penalised regression models were used to develop DNAm predictors for body mass index (BMI), smoking status, alcohol consumption, and educational attainment in a cohort of 5,100 individuals. Using an independent test cohort comprising 906 individuals, the proportion of phenotypic variance explained in each trait was examined for DNAm-based and genetic predictors. Receiver operator characteristic curves were generated to investigate the predictive performance of DNAm-based predictors, using dichotomised phenotypes. The relationship between DNAm scores and all-cause mortality (n = 214 events) was assessed via Cox proportional-hazards models.\n\nResultsThe DNAm-based predictors explained different proportions of the phenotypic variance for BMI (12%), smoking (60%), alcohol consumption (12%) and education (3%). The combined genetic and DNAm predictors explained 20% of the variance in BMI, 61% in smoking, 13% in alcohol consumption, and 6% in education. DNAm predictors for smoking, alcohol, and education but not BMI predicted mortality in univariate models. The predictors showed moderate discrimination of obesity (AUC=0.67) and alcohol consumption (AUC=0.75), and excellent discrimination of current smoking status (AUC=0.98). There was poorer discrimination of college-educated individuals (AUC=0.59).\n\nConclusionsDNAm predictors correlate with lifestyle factors that are associated with health and mortality. They may supplement DNAm-based predictors of age to identify the lifestyle profiles of individuals and predict disease risk.\n\nList of abbreviations

genomics

An epigenetic score for BMI based on DNA methylation correlates with poor physical health and major disease in the Lothian Birth Cohort 1936.

BackgroundThe relationship between obesity and adverse health is well established, but little is known about the contribution of DNA methylation to obesity-related health outcomes. Additionally, it is of interest whether such contributions are independent of those attributed by the most widely used clinical measure of body mass - the Body Mass Index (BMI).\n\nMethodWe tested whether an epigenetic BMI score accounts for inter-individual variation in health-related, cognitive, psychosocial and lifestyle outcomes in the Lothian Birth Cohort 1936 (n=903). Weights for the epigenetic BMI score were derived using penalised regression on methylation data from unrelated Generation Scotland participants (n=2566).\n\nResultsThe Epigenetic BMI score was associated with variables related to poor physical health (R2 ranges from 0.02-0.10), metabolic syndrome (R2 ranges from 0.01-0.09), lower crystallised intelligence (R2=0.01), lower health-related quality of life (R2=0.02), physical inactivity (R2=0.02), and social deprivation (R2=0.02). The epigenetic BMI score (per SD) was also associated with self-reported type 2 diabetes (OR 2.25, 95 % CI 1.74, 2.94), cardiovascular disease (OR 1.44, 95 % CI 1.23, 1.69) and high blood pressure (OR 1.21, 95% CI 1.13, 1.48; all at p<0.0011 after Bonferroni correction).\n\nConclusionsOur results show that regression models with epigenetic and phenotypic BMI scores as predictors account for a greater proportion of all outcome variables than either predictor alone, demonstrating independent and additive effects of epigenetic and phenotypic BMI scores.

epidemiology

DNA methylation age acceleration and risk factors for Alzheimer’s disease

INTRODUCTIONThe epigenetic clock is a DNA methylation-based estimate of biological age and is correlated with chronological age - the greatest risk factor for Alzheimers disease (AD). Genetic and environmental risk factors exist for AD, several of which are potentially modifiable. Here, we assess the relationship associations between the epigenetic clock and AD risk factors.\n\nMETHODSLinear mixed modelling was used to assess the relationship between age acceleration (the residual of biological age regressed onto chronological age) and AD risk factors relating to cognitive reserve, lifestyle, disease, and genetics in the Generation Scotland study (n=5,100).\n\nRESULTSWe report significant associations between the epigenetic clock and BMI, total:HDL cholesterol ratios, socioeconomic status, and smoking behaviour (Bonferroni-adjusted P<0.05).\n\nDISCUSSIONAssociations are present between environmental risk factors for AD and age acceleration. Measures to modify such risk factors might improve the risk profile for AD and the rate of biological ageing. Future longitudinal analyses are therefore warranted.

genomics

Genome-wide Meta-analysis of 158,000 Individuals of European Ancestry Identifies Three Loci Associated with Chronic Back Pain

OBJECTIVESTo conduct a genome-wide association study (GWAS) meta-analysis of chronic back pain (CBP).\n\nMETHODSAdults of European ancestry were included from 16 cohorts in Europe and North America. CBP cases were defined as those reporting back pain present for >3-6 months; non-cases were included as comparisons (\"controls\"). Each cohort conducted genotyping using commercially available arrays followed by imputation. GWAS used logistic regression models with additive genetic effects, adjusting for age, sex, study-specific covariates, and population substructure. The threshold for genome-wide significance in the fixed-effect inverse-variance weighted meta-analysis was p<5x10-8. Suggestive (p<5x10-7) and genome-wide significant (p<5x10-8) variants were carried forward for replication or further investigation in an independent sample.\n\nRESULTSThe discovery sample was comprised of 158,025 individuals, including 29,531 CBP cases. A genome-wide significant association was found for the intronic variant rs12310519 in SOX5 (OR 1.08, p=7.2x10-10). This was subsequently replicated in an independent sample of 283,752 subjects, including 50,915 cases (OR 1.06, p=5.3x10-11), and exceeded genome-wide significance in joint meta-analysis (0R=1.07, p=4.5x10-19). We found suggestive associations at three other loci in the discovery sample, two of which exceeded genome-wide significance in joint meta-analysis: an intergenic variant, rs7833174, located between CCDC26 and GSDMC (OR 1.05, p=4.4x10-13), and an intronic variant, rs4384683, in DCC (OR 0.97, p=2.4x10-10).\n\nDISCUSSIONIn this first reported meta-analysis of GWAS for CBP, we identified and replicated a genetic locus associated with CBP (SOX5). We also identified 2 other loci that reached genome-wide significance in a 2-stage joint meta-analysis (CCDC26/GSDMC and DCC).

genomics

Genetic and environmental determinants of stressful life events and their overlap with depression and neuroticism

BackgroundStressful life events (SLEs) and neuroticism are risk factors for major depressive disorder (MDD). However, SLEs and neuroticism are heritable traits that are correlated with genetic risk for MDD. In the current study, we sought to investigate the genetic and environmental contributions to SLEs in a large family-based sample, and quantify any genetic overlap with MDD and neuroticism.\n\nMethodsA subset of Generation Scotland: the Scottish Family Health Study, consisting of 9618 individuals comprise the present study. We estimated the heritability of SLEs using pedigree-based and molecular genetic data. The environment was assessed by modelling familial, couple and sibling components. Using polygenic risk scores (PRS) and LD score regression we analysed the genetic overlap between MDD, neuroticism and SLEs.\n\nResultsPast 6-month life events were positively correlated with lifetime MDD status ({beta}=0.21, r2=1.1%, p=2.5 x 10-25) and neuroticism ({beta} =0.13, r2=1.9%, p=1.04 x 10-37). Common SNPs explained 8% of the variance in personal life events (those directly affecting the individual) (S.E.=0.03, p=9 x 10-4). A significant effect of couple environment accounted for 13% (S.E.=0.03, p=0.016) of variation in SLEs. PRS analyses found that individuals with higher PRS for MDD reported more SLEs ({beta} =0.05, r2=0.3%, p=3 x 10-5). LD score regression demonstrated genetic correlations between MDD and both SLEs (rG=0.33, S.E.=0.08) and neuroticism (rG=0.15, S.E.=0.07).\n\nConclusionsThese findings suggest that SLEs are partially heritable and this heritability is shared with risk for MDD and neuroticism. Further work should determine the causal direction and source of these associations.

genetics

Data Resource Profile: Generation Scotland Electronic Health Record

This paper provides the first detailed demonstration of the research value of the Electronic Health Record (EHR) linked to research data in Generation Scotland Scottish Family Health Study (GS:SFHS) participants, together with how to access this data. The structured, coded variables in the routine biochemistry, prescribing and morbidity records in particular represent highly valuable phenotypic data for a genomics research resource. Access to a wealth of other specialized datasets including cancer, mental health and maternity inpatient information is also possible through the same straightforward and transparent application process. The Electronic Health Record linked dataset is a key component of GS:SFHS, a biobank conceived in 1999 for the purpose of studying the genetics of health areas of current and projected public health importance. Over 24,000 adults were recruited from 2006 to 2011, with broad and enduring written informed consent for biomedical research. Consent was obtained from 23,603 participants for GS:SFHS study data to be linked to their Scottish National Health Service (NHS) records, using their Community Health Index (CHI) number. This identifying number is used for NHS Scotland procedures (registrations, attendances, samples, prescribing and investigations) and allows healthcare records for individuals to be linked across time and location. Here, we describe the NHS EHR dataset on the sub-cohort of 20,032 GS:SFHS participants with consent and mechanism for record linkage plus extensive genetic data. Together with existing study phenotypes, including family history and environmental exposures such as smoking, the EHR is a rich resource of real world data that can be used in research to characterise the health trajectory of participants, available at low cost and a high degree of timeliness, matched to DNA, urine and serum samples and genome-wide genetic information.

genetics

Polymorphisms in the vitamin D receptor gene are associated with reduced rate of sputum culture conversion in multidrug-resistant tuberculosis patients in South Africa

BackgroundVitamin D modulates the inflammatory and immune response to tuberculosis (TB) and also mediates the induction of the antimicrobial peptide cathelicidin. Deficiency of 25-hydroxyvitamin D and single nucleotide polymorphisms (SNPs) in the vitamin D receptor (VDR) gene may increase the risk of TB disease and decrease culture conversion rates in drug susceptible TB. Whether these VDR SNPs are found in African populations or impact multidrug-resistant (MDR) TB treatment has not been established. We aimed to determine if SNPs in the VDR gene were associated with sputum culture conversion among a cohort of MDR TB patients in South Africa.\n\nMethodsWe conducted a prospective cohort study of adult MDR TB patients receiving second-line TB treatment in KwaZulu-Natal province. Subjects had monthly sputum cultures performed. In a subset of participants, whole blood samples were obtained for genomic analyses. Genomic DNA was extracted and genotyped with Affymetrix Axiom Pan-African Array. Cox proportional models were used to determine the association between VDR SNPs and rate of culture conversion.\n\nResultsGenomic analyses were performed on 91 MDR TB subjects enrolled in the sub-study; 60% were female and median age was 35 years (interquartile range [IQR] 29-42). Smoking was reported by 21% of subjects and most subjects had HIV (80%), were smear negative (57%), and had cavitary disease (55%). Overall, 87 (96%) subjects initially converted cultures to negative, with median time to culture conversion of 57 days (IQR 17-114). Of 121 VDR SNPs examined, 10 were significantly associated (p<0.01) with rate of sputum conversion in multivariable analyses. Each additional risk allele on SNP rs74085240 delayed culture conversion significantly (adjusted hazard ratio 0.30, 95% confidence interval 0.14-0.67).\n\nConclusionsPolymorphisms in the VDR gene were associated with rate of sputum culture conversion in MDR TB patients in this high HIV prevalence setting in South Africa.\n\nAuthor contributionsMJM, YVS, JCMB, SS, and NRG conceived and designed the study and drafted the initial manuscript. MJM, YVS, YN, and QH performed the data analyses. All authors contributed to interpretation of the data, revised the manuscript, and approved the final version.\n\nThe findings and conclusions in this article are those of the authors and do not necessarily represent the views of the Centers for Disease Control and Prevention (CDC). The use of trade names and commercial sources is for identification only and does not imply endorsement by the CDC.

epidemiology

Spatial distribution of extensively drug-resistant tuberculosis (XDR-TB) patients in KwaZulu-Natal, South Africa

BackgroundKwaZulu-Natal province, South Africa, has among the highest burden of XDR-TB worldwide with the majority of cases occurring due to transmission. Poor access to health facilities can be a barrier to timely diagnosis and treatment of TB, which can contribute to ongoing transmission. We sought to determine the geographic distribution of XDR-TB patients and proximity to health facilities in KwaZulu-Natal.\n\nMethodsWe recruited adults and children with XDR-TB diagnosed in KwaZulu-Natal. We calculated distance and time from participants home to the closest hospital or clinic, as well as to the actual facility that diagnosed XDR-TB, using tools within ArcGIS Network analyst. Speed of travel was assigned to road classes based on Department of Transport regulations. Results were compared to guidelines for the provision of social facilities in South Africa: 5km to a clinic and 30km to a hospital.\n\nResultsDuring 2011-2014, 1027 new XDR-TB cases were diagnosed throughout all 11 districts of KwaZulu-Natal, of whom 404 (39%) were enrolled and had geospatial data collected. Participants would have had to travel a mean distance of 2.9 km (CI 95%: 1.8-4.1) to the nearest clinic and 17.6 km (CI 95%: 11.4-23.8) to the nearest hospital. Actual distances that participants travelled to the health facility that diagnosed XDR-TB ranged from <10 km (n=143, 36%) to >50 km (n=109, 27%). The majority (77%) of participants travelled farther than the recommended distance to a clinic (5 km) and 39% travelled farther than the recommended distance to a hospital (30 km). Nearly half (46%) of participants were diagnosed at a health facility in eThekwini district, of whom, 36% resided outside the Durban metropolitan area.\n\nConclusionsXDR-TB cases are widely distributed throughout KwaZulu-Natal province with a denser focus in eThekwini district. Patients travelled long distances to the health facility where they were diagnosed with XDR-TB, suggesting a potential role for migration or transportation in the XDR-TB epidemic.

epidemiology

Genome-wide association study of alcohol consumption and genetic overlap with other health-related traits in UK Biobank (N=112,117).

Alcohol consumption has been linked to over 200 diseases and is responsible for over 5% of the global disease burden. Well known genetic variants in alcohol metabolizing genes, e.g. ALDH2, ADH1B, are strongly associated with alcohol consumption but have limited impact in European populations where they are found at low frequency. We performed a genome-wide association study (GWAS) of self-reported alcohol consumption in 112,117 individuals in the UK Biobank (UKB) sample of white British individuals. We report significant genome-wide associations at 8 independent loci. These include SNPs in alcohol metabolizing genes (ADH1B/ADH1C/ADH5) and 2 loci in KLB, a gene recently associated with alcohol consumption. We also identify SNPs at novel loci including GCKR, PXDN, CADM2 and TNFRSF11A. Gene-based analyses found significant associations with genes implicated in the neurobiology of substance use (CRHR1, DRD2), and genes previously associated with alcohol consumption (AUTS2). GCTA-GREML analyses found a significant SNP-based heritability of self-reported alcohol consumption of 13% (S.E.=0.01). Sex-specific analyses found largely overlapping GWAS loci and the genetic correlation between male and female alcohol consumption was 0.73 (S.E.=0.09, p-value = 1.37 x 10-16). Using LD score regression, genetic overlap was found between alcohol consumption and schizophrenia (rG=0.13, S.E=0.04), HDL cholesterol (rG=0.21, S.E=0.05), smoking (rG=0.49, S.E=0.06) and various anthropometric traits (e.g. Overweight, rG=-0.19, S.E.=0.05). This study replicates the association between alcohol consumption and alcohol metabolizing genes and KLB, and identifies 4 novel gene associations that should be the focus of future studies investigating the neurobiology of alcohol consumption.

genetics