Search bioRxivSearch

Biology subjects

Mahajan, A.

Publications and source records attributed to Mahajan, A..

15 recordsLinked to original sources

Maternal and fetal genetic effects on birth weight and their relevance to cardio-metabolic risk factors

Birth weight (BW) variation is influenced by fetal and maternal genetic and non-genetic factors, and has been reproducibly associated with future cardio-metabolic health outcomes. These associations have been proposed to reflect the lifelong consequences of an adverse intrauterine environment. In earlier work, we demonstrated that much of the negative correlation between BW and adult cardio-metabolic traits could instead be attributable to shared genetic effects. However, that work and other previous studies did not systematically distinguish the direct effects of an individuals own genotype on BW and subsequent disease risk from indirect effects of their mothers correlated genotype, mediated by the intrauterine environment. Here, we describe expanded genome-wide association analyses of own BW (n=321,223) and offspring BW (n=230,069 mothers), which identified 278 independent association signals influencing BW (214 novel). We used structural equation modelling to decompose the contributions of direct fetal and indirect maternal genetic influences on BW, implicating fetal- and maternal-specific mechanisms. We used Mendelian randomization to explore the causal relationships between factors influencing BW through fetal or maternal routes, for example, glycemic traits and blood pressure. Direct fetal genotype effects dominate the shared genetic contribution to the association between lower BW and higher type 2 diabetes risk, whereas the relationship between lower BW and higher later blood pressure (BP) is driven by a combination of indirect maternal and direct fetal genetic effects: indirect effects of maternal BP-raising genotypes act to reduce offspring BW, but only direct fetal genotype effects (once inherited) increase the offsprings later BP. Instrumental variable analysis using maternal BW-lowering genotypes to proxy for an adverse intrauterine environment provided no evidence that it causally raises offspring BP. In successfully separating fetal from maternal genetic effects, this work represents an important advance in genetic studies of perinatal outcomes, and shows that the association between lower BW and higher adult BP is attributable to genetic effects, and not to intrauterine programming.

genetics

Variation in the plasma membrane monoamine transporter (PMAT, encoded in SLC29A4) and organic cation transporter 1 (OCT1, encoded in SLC22A1) and gastrointestinal intolerance to metformin in type 2 diabetes: an IMI DIRECT study

Objectives20-30% of patients with metformin treated type 2 diabetes experience gastrointestinal side effects leading to premature discontinuation in 5-10% of the cases. Gastrointestinal intolerance may reflect localised high concentrations of metformin in the gut. We hypothesized that reduced transport of metformin into the circulation via the plasma membrane monoamine transporter (PMAT) and organic cation transporter 1 (OCT1) could increase the risk of severe GI side effects.\n\nResearch Design and MethodsThe study included 286 severe metformin intolerant and 1128 tolerant individuals from the IMI DIRECT consortium. We assessed the association of patient characteristics, concomitant medication and the burden of mutations in the SLC29A4 and SLC22A1, genes that encode PMAT and OCT1, respectively, on odds of metformin intolerance using a logistic regression model.\n\nResultsWomen (p < 0.001) and older people (p < 0.001) were more likely to develop metformin intolerance. Concomitant use of metformin transporter inhibiting drugs increased the odds of intolerance by more than 70% (OR = 1.72 [1.26-2.32], p < 0.001). In a logistic regression model adjusted for age, sex, weight and population substructure, the G allele at rs3889348 (SLC29A4) was associated with GI intolerance (OR = 1.34[1.09-1.65], p = 0.005). rs3889348 is the top cis-eQTL for SLC29A4 in gut tissue where carriers of the G allele had reduced expression. Homozygous carriers of the G allele treated with metformin transporter inhibiting drugs had over three times higher odds of intolerance compared to carriers of no G allele and not treated with inhibiting drugs (OR = 3.23 [1.71-6.39], p < 0.001). Using a genetic risk score (GRS) derived from SLC29A4 (rs3889348) and previously reported SLC22A1 variants (M420del, R61C, G401S), the odds of intolerance was more than twice in individuals who carry three or more risk alleles compared with those carrying none (OR = 2.15 [1.20-4.12], p = 0.01).\n\nConclusionsThese results suggest that intestinal metformin transporters and concomitant medications play an important role in gastrointestinal side effects of metformin.

genetics

Variants in the fetal genome near pro-inflammatory cytokine genes on 2q13 are associated with gestational duration

The duration of pregnancy is influenced by fetal and maternal genetic and non-genetic factors. We conducted a fetal genome-wide association meta-analysis of gestational duration, and early preterm, preterm, and postterm birth in 84,689 infants. One locus on chromosome 2q13 was associated with gestational duration; the association was replicated in 9,291 additional infants (combined P = 3.96 x 10-14). Analysis of 15,536 mother-child pairs showed that the association was driven by fetal rather than maternal genotype. Functional experiments showed that the lead SNP, rs7594852, alters the binding of the HIC1 transcriptional repressor. Genes at the locus include several interleukin 1 family members with roles in pro-inflammatory pathways that are central to the process of parturition. Further understanding of the underlying mechanisms will be of great public health importance, since giving birth either before or after the window of term gestation is associated with increased morbidity and mortality.

genetics

Trans-ethnic genome-wide association study provides insight into effector genes and molecular mechanisms for kidney function and highlights a causal effect on kidney-specific disease aetiologies

Chronic kidney disease (CKD) affects [~]10% of the global population, with considerable ethnic differences in prevalence and aetiology. We assembled genome-wide association studies (GWAS)1-3 of estimated glomerular filtration rate (eGFR), a measure of kidney function that defines CKD, in 312,468 individuals from four ancestry groups. We identified 93 loci (20 novel), which were delineated to 127 distinct association signals. These signals were homogenous across ancestries, and were enriched for protein-coding exons, kidney-specific histone modifications, and transcription factor binding sites for HDAC2 and EZH2. Fine-mapping revealed 40 high-confidence variants driving eGFR associations and highlighted potential causal genes with cell-type specific expression in glomerulus, and proximal and distal nephron. Mendelian randomisation (MR) supported causal effects of eGFR on overall and cause-specific CKD, kidney stone formation, diastolic blood pressure (DBP) and hypertension. These results define novel molecular mechanisms and effector genes for eGFR, offering insight into clinical outcomes and routes to CKD treatment development.

genetics

Pituitary Carcinoma: The University of Texas MD Anderson Cancer Center Experience.

Background: Pituitary carcinoma (PC) is an aggressive neuroendocrine tumor diagnosed when a pituitary adenoma (PA) becomes metastatic. PCs are typically resistant to therapy and frequently recur. Recently, treatment with temozolomide (TMZ) has shown promising results, although the lack of prospective trials limits accurate assessment. Methods: We describe a single-center experience in managing PC over a 22-year period and review previously published PC series. Results: 17 patients were identified. Median age at PC diagnosis was 44 years (range 16-82), and the median PA-to-PC conversion time was 5 years (range 1-29). Median follow-up was 28 months (range 8-158) with 7 deaths. Most PC were hormone-positive based on immunohistochemistry (n=12): ACTH (n=5), PRL (n=4), LH/FSH (n=2), GH (n=1). All patients underwent at least one resection and one course of radiation after PC diagnosis. Immunohistochemistry showed high Ki-67 labeling index (>3%) in 10/15 cases. Eight patients (47%) had metastases only to the CNS, and 6 (35%) had combined CNS and systemic metastases. The most commonly used chemotherapy was TMZ, and TMZ-based therapy was associated with the longest period of disease control in 12 (71%) cases, as well as the longest period from PC diagnosis to first progression in 8 (47%) cases. The 2, 3 and 5-year survival rate of the entire cohort was 71%, 59% and 35%, respectively. All patients surviving >5 years were treated with TMZ-based therapy. Conclusions: PC treatment requires a multidisciplinary approach and multimodality therapy including surgery, radiation and chemotherapy. TMZ-based therapy was associated with higher survival rates and longer disease control.\n\nPrecisWe describe 17 PC patients who were diagnosed and treated at MDACC over a 22-year period. We have found that TMZ-based therapy correlated with longer disease control and higher survival rate.

cancer biology

Genetic discovery and translational decision support from exome sequencing of 20,791 type 2 diabetes cases and 24,440 controls from five ancestries

Protein-coding genetic variants that strongly affect disease risk can provide important clues into disease pathogenesis. Here we report an exome sequence analysis of 20,791 type 2 diabetes (T2D) cases and 24,440 controls from five ancestries. We identify rare (minor allele frequency<0.5%) variant gene-level associations in (a) three genes at exome-wide significance, including a T2D-protective series of >30 SLC30A8 alleles, and (b) within 12 gene sets, including those corresponding to T2D drug targets (p=6.1x10-3) and candidate genes from knockout mice (p=5.2x10-3). Within our study, the strongest T2D rare variant gene-level signals explain at most 25% of the heritability of the strongest common single-variant signals, and the rare variant gene-level effect sizes we observe in established T2D drug targets will require 110K-180K sequenced cases to exceed exome-wide significance. To help prioritize genes using associations from current smaller sample sizes, we present a Bayesian framework to recalibrate association p-values as posterior probabilities of association, estimating that reaching p<0.05 (p<0.005) in our study increases the odds of causal T2D association for a nonsynonymous variant by a factor of 1.8 (5.3). To help guide target or gene prioritization efforts, our data are freely available for analysis at www.type2diabetesgenetics.org.

genetics

PROTEIN-CODING VARIANTS IMPLICATE NOVEL GENES RELATED TO LIPID HOMEOSTASIS CONTRIBUTING TO BODY FAT DISTRIBUTION

Body fat distribution is a heritable risk factor for a range of adverse health consequences, including hyperlipidemia and type 2 diabetes. To identify protein-coding variants associated with body fat distribution, assessed by waist-to-hip ratio adjusted for body mass index, we analyzed 228,985 predicted coding and splice site variants available on exome arrays in up to 344,369 individuals from five major ancestries for discovery and 132,177 independent European-ancestry individuals for validation. We identified 15 common (minor allele frequency, MAF[&ge;]5%) and 9 low frequency or rare (MAF<5%) coding variants that have not been reported previously. Pathway/gene set enrichment analyses of all associated variants highlight lipid particle, adiponectin level, abnormal white adipose tissue physiology, and bone development and morphology as processes affecting fat distribution and body shape. Furthermore, the cross-trait associations and the analyses of variant and gene function highlight a strong connection to lipids, cardiovascular traits, and type 2 diabetes. In functional follow-up analyses, specifically in Drosophila RNAi-knockdown crosses, we observed a significant increase in the total body triglyceride levels for two genes (DNAH10 and PLXND1). By examining variants often poorly tagged or entirely missed by genome-wide association studies, we implicate novel genes in fat distribution, stressing the importance of interrogating low-frequency and protein-coding variants.

genetics

Developing a network view of type 2 diabetes risk pathways through integration of genetic, genomic and functional data

Genome wide association studies (GWAS) have identified several hundred susceptibility loci for Type 2 Diabetes (T2D). One critical, but unresolved, issue concerns the extent to which the mechanisms through which these diverse signals influencing T2D predisposition converge on a limited set of biological processes. However, the causal variants identified by GWAS mostly fall into non-coding sequence, complicating the task of defining the effector transcripts through which they operate. Here, we describe implementation of an analytical pipeline to address this question. First, we integrate multiple sources of genetic, genomic, and biological data to assign positional candidacy scores to the genes that map to T2D GWAS signals. Second, we introduce genes with high scores as seeds within a network optimization algorithm (the asymmetric prize-collecting Steiner Tree approach) which uses external, experimentally-confirmed protein-protein interaction (PPI) data to generate high confidence subnetworks. Third, we use GWAS data to test the T2D-association enrichment of the \"non-seed\" proteins introduced into the network, as a measure of the overall functional connectivity of the network. We find: (a) non-seed proteins in the T2D protein-interaction network so generated (comprising 705 nodes) are enriched for association to T2D (p=0.0014) but not control traits; (b) stronger T2D-enrichment for islets than other tissues when we use RNA expression data to generate tissue-specific PPI networks; and (c) enhanced enrichment (p=3.9xl0-5) when we combine analysis of the islet-specific PPI network with a focus on the subset of T2D GWAS loci which act through defective insulin secretion. These analyses reveal a pattern of non-random functional connectivity between causal candidate genes atT2D GWAS loci, and highlight the products of genes including YWHAG, SMAD4 or CDK2 as contributors to T2D-relevant islet dysfunction. The approach we describe can be applied to other complex genetic and genomic data sets, facilitating integration of diverse data types into disease-associated networks.\n\nAuthor summaryWe were interested in the following question: as we discover more and more genetic variants associated with a complex disease, such as type 2 diabetes, will the biological pathways implicated by those variants proliferate, or will the biology converge onto a more limited set of aetiological processes? To address this, we first took the 1895 genes that map to ~100 type 2 diabetes association signals, and pruned these to a set of 451 for which combined genetic, genomic and biological evidence assigned the strongest candidacy with respect to type 2 diabetes pathogenesis. We then sought to maximally connect these genes within a curated protein-protein interaction network. We found that proteins brought into the resulting diabetes interaction network were themselves enriched for diabetes association signals as compared to appropriate control proteins. Furthermore, when we used tissue-specific RNA abundance data to filter the generic protein-protein network, we found that the enrichment for type 2 diabetes association signals was enhanced within a network filtered for pancreatic islet expression, particularly when we selected the subset of diabetes association signals acting through reduced insulin secretion. Our data demonstrate convergence of the biological processes involved in type 2 diabetes pathogenesis and highlight novel contributors.

bioinformatics

Preoperative predictions of in-hospital mortality using electronic medical record data

BackgroundPredicting preoperative in-hospital mortality using readily-available electronic medical record (EMR) data can aid clinicians in accurately and rapidly determining surgical risk. While previous work has shown that the American Society of Anesthesiologists (ASA) Physical Status Classification is a useful, though subjective, feature for predicting surgical outcomes, obtaining this classification requires a clinician to review the patients medical records. Our goal here is to create an improved risk score using electronic medical records and demonstrate its utility in predicting in-hospital mortality without requiring clinician-derived ASA scores.\n\nMethodsData from 49,513 surgical patients were used to train logistic regression, random forest, and gradient boosted tree classifiers for predicting in-hospital mortality. The features used are readily available before surgery from EMR databases. A gradient boosted tree regression model was trained to impute the ASA Physical Status Classification, and this new, imputed score was included as an additional feature to preoperatively predict in-hospital post-surgical mortality. The preoperative risk prediction was then used as an input feature to a deep neural network (DNN), along with intraoperative features, to predict postoperative in-hospital mortality risk. Performance was measured using the area under the receiver operating characteristic (ROC) curve (AUC).\n\nResultsWe found that the random forest classifier (AUC 0.921, 95%CI 0.908-0.934) outperforms logistic regression (AUC 0.871, 95%CI 0.841-0.900) and gradient boosted trees (AUC 0.897, 95%CI 0.881-0.912) in predicting in-hospital post-surgical mortality. Using logistic regression, the ASA Physical Status Classification score alone had an AUC of 0.865 (95%CI 0.848-0.882). Adding preoperative features to the ASA Physical Status Classification improved the random forest AUC to 0.929 (95%CI 0.915-0.943). Using only automatically obtained preoperative features with no clinician intervention, we found that the random forest model achieved an AUC of 0.921 (95%CI 0.908-0.934). Integrating the preoperative risk prediction into the DNN for postoperative risk prediction results in an AUC of 0.924 (95%CI 0.905-0.941), and with both a preoperative and postoperative risk score for each patient, we were able to show that the mortality risk changes over time.\n\nConclusionsFeatures easily extracted from EMR data can be used to preoperatively predict the risk of in-hospital post-surgical mortality in a fully automated fashion, with accuracy comparable to models trained on features that require clinical expertise. This preoperative risk score can then be compared to the postoperative risk score to show that the risk changes, and therefore should be monitored longitudinally over time.\n\nAuthor summaryRapid, preoperative identification of those patients at highest risk for medical complications is necessary to ensure that limited infrastructure and human resources are directed towards those most likely to benefit. Existing risk scores either lack specificity at the patient level, or utilize the American Society of Anesthesiologists (ASA) physical status classification, which requires a clinician to review the chart. In this manuscript we report on using machine-learning algorithms, specifically random forest, to create a fully automated score that predicts preoperative in-hospital mortality based solely on structured data available at the time of surgery. This score has a higher AUC than both the ASA physical status score and the Charlson comorbidity score. Additionally, we integrate this score with a previously published postoperative score to demonstrate the extent to which patient risk changes during the perioperative period.

bioinformatics

Discovery of biomarkers for glycaemic deterioration before and after the onset of type 2 diabetes: an overview of the data from the epidemiological studies within the IMI DIRECT Consortium

Abstract/SummaryO_ST_ABSBackground and aimsC_ST_ABSUnderstanding the aetiology, clinical presentation and prognosis of type 2 diabetes (T2D) and optimizing its treatment might be facilitated by biomarkers that help predict a persons susceptibility to the risk factors that cause diabetes or its complications, or response to treatment. The IMI DIRECT (Diabetes Research on Patient Stratification) Study is a European Union (EU) Innovative Medicines Initiative (IMI) project that seeks to test these hypotheses in two recently established epidemiological cohorts. Here, we describe the characteristics of these cohorts at baseline and at the first main follow-up examination (18-months).\n\nMaterials and methodsFrom a sampling-frame of 24,682 European-ancestry adults in whom detailed health information was available, participants at varying risk of glycaemic deterioration were identified using a risk prediction algorithm and enrolled into a prospective cohort study (n=2127) undertaken at four study centres across Europe (Cohort 1: prediabetes). We also recruited people from clinical registries with recently diagnosed T2D (n=789) into a second cohort study (Cohort 2: diabetes). The two cohorts were studied in parallel with matched protocols. Endogenous insulin secretion and insulin sensitivity were modelled from frequently sampled 75g oral glucose tolerance (OGTT) in Cohort 1 and with mixed-meal tolerance tests (MMTT) in Cohort 2. Additional metabolic biochemistry was determined using blood samples taken when fasted and during the tolerance tests. Body composition was assessed using MRI and lifestyle measures through self-report and objective methods.\n\nResultsUsing ADA-2011 glycaemic categories, 33% (n=693) of Cohort 1 (prediabetes) had normal glucose regulation (NGR), and 67% (n=1419) had impaired glucose regulation (IGR). 76% of the cohort was male, age=62(6.2) years; BMI=27.9(4.0) kg/m2; fasting glucose=5.7(0.6) mmol/l; 2-hr glucose=5.9(1.6) mmol/l [mean(SD)]. At follow-up, 18.6(1.4) months after baseline, fasting glucose=5.8(0.6) mmol/l; 2-hr OGTT glucose=6.1(1.7) mmol/l [mean(SD)]. In Cohort 2 (diabetes): 65% (n=508) were lifestyle treated (LS) and 35% (n=271) were lifestyle + metformin treated (LS+MET). 58% of the cohort was male, age=62(8.1) years; BMI=30.5(5.0) kg/m2; fasting glucose=7.2(1.4)mmol/l; 2-hr glucose=8.6(2.8) mmol/l [mean(SD)]. At follow-up, 18.2(0.6) months after baseline, fasting glucose=7.8(1.8) mmol/l; 2-hr MMTT glucose=9.5(3.3) mmol/l [mean(SD)].\n\nConclusionThe epidemiological IMI DIRECT cohorts are the most intensely characterised prospective studies of glycaemic deterioration to date. Data from these cohorts help illustrate the heterogeneous characteristics of people at risk of or with T2D, highlighting the rationale for biomarker stratification of the disease - the primary objective of the IMI DIRECT consortium.\n\nAbbreviations

epidemiology

Large-Scale Genome-Wide Meta Analysis of Polycystic Ovary Syndrome Suggests Shared Genetic Architecture for Different Diagnosis Criteria.

Polycystic ovary syndrome (PCOS) is a disorder characterized by hyperandrogenism, ovulatory dysfunction and polycystic ovarian morphology. Affected women frequently have metabolic disturbances including insulin resistance and dysregulation of glucose homeostasis. PCOS is diagnosed with two different sets of diagnostic criteria, resulting in a phenotypic spectrum of PCOS cases. The genetic similarities between cases diagnosed with different criteria have been largely unknown. Previous studies in Chinese and European subjects have identified 16 loci associated with risk of PCOS. We report a meta-analysis from 10,074 PCOS cases and 103,164 controls of European ancestry and characterisation of PCOS related traits. We identified 3 novel loci (near PLGRKT, ZBTB16 and MAPRE1), and provide replication of 11 previously reported loci. Identified variants were associated with hyperandrogenism, gonadotropin regulation and testosterone levels in affected women. Genetic correlations with obesity, fasting insulin, type 2 diabetes, lipid levels and coronary artery disease indicate shared genetic architecture between metabolic traits and PCOS. Mendelian randomization analyses suggested variants associated with body mass index, fasting insulin, menopause timing, depression and male-pattern balding play a causal role in PCOS. Only one locus differed in its association by diagnostic criteria, otherwise the genetic architecture was similar between PCOS diagnosed by self-report and PCOS diagnosed by NIH or Rotterdam criteria across common variants at 13 loci.

genetics

Fine-mapping of an expanded set of type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps

We aggregated genome-wide genotyping data from 32 European-descent GWAS (74,124 T2D cases, 824,006 controls) imputed to high-density reference panels of >30,000 sequenced haplotypes. Analysis of {small tilde}27M variants ({small tilde}21M with minor allele frequency [MAF]<5%), identified 243 genome-wide significant loci (p<5x10-8; MAF 0.02%-50%; odds ratio [OR] 1.04-8.05), 135 not previously-implicated in T2D-predisposition. Conditional analyses revealed 160 additional distinct association signals (p<10-5) within the identified loci. The combined set of 403 T2D-risk signals includes 56 low-frequency (0.5%[&le;]MAF<5%) and 24 rare (MAF<0.5%) index SNPs at 60 loci, including 14 with estimated allelic OR>2. Forty-one of the signals displayed effect-size heterogeneity between BMI-unadjusted and adjusted analyses. Increased sample size and improved imputation led to substantially more precise localisation of causal variants than previously attained: at 51 signals, the lead variant after fine-mapping accounted for >80% posterior probability of association (PPA) and at 18 of these, PPA exceeded 99%. Integration with islet regulatory annotations enriched for T2D association further reduced median credible set size (from 42 variants to 32) and extended the number of index variants with PPA>80% to 73. Although most signals mapped to regulatory sequence, we identified 18 genes as human validated therapeutic targets through coding variants that are causal for disease. Genome wide chip heritability accounted for 18% of T2D-risk, and individuals in the 2.5% extremes of a polygenic risk score generated from the GWAS data differed >9-fold in risk. Our observations highlight how increases in sample size and variant diversity deliver enhanced discovery and single-variant resolution of causal T2D-risk alleles, and the consequent impact on mechanistic insights and clinical translation.

genomics

Integration of human pancreatic islet genomic data refines regulatory mechanisms at Type 2 Diabetes susceptibility loci

Human genetic studies have emphasised the dominant contribution of pancreatic islet dysfunction to development of Type 2 Diabetes (T2D). However, limited annotation of the islet epigenome has constrained efforts to define the molecular mechanisms mediating the, largely regulatory, signals revealed by Genome-Wide Association Studies (GWAS). We characterised patterns of chromatin accessibility (ATAC-seq, n=17) and DNA methylation (whole-genome bisulphite sequencing, n=10) in human islets, generating high-resolution chromatin state maps through integration with established ChIP-seq marks. We found enrichment of GWAS signals for T2D and fasting glucose was concentrated in subsets of islet enhancers characterised by open chromatin and hypomethylation, with the former annotation predominant. At several loci (including CDC123, ADCY5, KLHDC5) the combination of fine-mapping genetic data and chromatin state enrichment maps, supplemented by allelic imbalance in chromatin accessibility pinpointed likely causal variants. The combination of increasingly-precise genetic and islet epigenomic information accelerates definition of causal mechanisms implicated in T2D pathogenesis.

genomics

Meta-analysis of exome array data identifies six novel genetic loci for lung function

Over 90 regions of the genome have been associated with lung function to date, many of which have also been implicated in chronic obstructive pulmonary disease (COPD). We carried out meta-analyses of exome array data and three lung function measures: forced expiratory volume in one second (FEV1), forced vital capacity (FVC) and the ratio of FEV1 to FVC (FEV1/FVC). These analyses by the SpiroMeta and CHARGE consortia included 60,749 individuals of European ancestry from 23 studies, and 7,721 individuals of African Ancestry from 5 studies in the discovery stage, with follow-up in up to 111,556 independent individuals. We identified significant (P<2{middle dot}8x10-7) associations with six SNPs: a nonsynonymous variant in RPAP1, which is predicted to be damaging, three intronic SNPs (SEC24C, CASC17 and UQCC1) and two intergenic SNPs near to LY86 and FGF10. eQTL analyses found evidence for regulation of gene expression at three signals and implicated several genes including TYRO3 and PLAU. Further interrogation of these loci could provide greater understanding of the determinants of lung function and pulmonary disease.

genetics

Refining The Accuracy Of Validated Target Identification Through Coding Variant Fine-Mapping In Type 2 Diabetes

Identification of coding variant associations for complex diseases offers a direct route to biological insight, but is dependent on appropriate inference concerning the causal impact of those variants on disease risk. We aggregated exome-array and exome sequencing data for 81,412 type 2 diabetes (T2D) cases and 370,832 controls of diverse ancestry, identifying 40 distinct coding variant association signals (at 38 loci) reaching significance (p<2.2x10-7). Of these, 16 represent novel associations mapping outside known genome-wide association study (GWAS) signals. We make two important observations. First, despite a threefold increase in sample size over previous efforts, only five of the 40 signals are driven by variants with minor allele frequency <5%, and we find no evidence for low-frequency variants with allelic odds ratio >1.36. Second, we used GWAS data from 50,160 T2D cases and 465,272 controls to fine-map associated coding variants in their regional context, with and without additional weighting, to account for the global enrichment of complex trait association signals in coding exons. We demonstrate convincing support (posterior probability >80% under the \"annotation-weighted\" model) that coding variants are causal for the association at 16 of the 40 signals (including novel signals involving POC5 p.His36Arg, ANKH p.Arg187Gln, WSCD2 p.Thr113Ile, PLCB3 p.Ser778Leu, and PNPLA3 p.Ile148Met). However, one third of coding variant association signals represent \"false leads\" at which naive analysis would have led to an erroneous inference regarding the effector transcript mediating the signal. Accurate identification of validated targets is dependent on correct specification of the contribution of coding and non-coding mediated mechanisms at associated loci.

genetics