Search bioRxivSearch

Biology subjects

Gagliano Taliun, S. A.

Publications and source records attributed to Gagliano Taliun, S. A..

5 recordsLinked to original sources

ERASE: Extended Randomization for assessment of annotation enrichment in ASE datasets

Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with various human phenotypes and many of these loci are thought to act at a molecular level by regulating gene expression. Detection of allele specific expression (ASE), namely preferential usage of an allele at a transcribed locus, is an increasingly important means of studying the genetic regulation of gene expression. However, there are currently a paucity of tools available to link ASE sites with GWAS risk loci. Existing integration methods first use ASE sites to infer cis-acting expression quantitative trait loci (eQTL) and then apply eQTL-based approaches. ERASE is a method that assesses the enrichment of risk loci amongst ASE sites directly. Furthermore, ERASE enables additional biological insights to be made through the addition of other SNP level annotations. ERASE is based on a randomization approach and controls for read depth, a significant confounder in ASE analyses. In this paper, we demonstrate that ERASE can efficiently detect the enrichment of eQTLs and risk loci within ASE data and that it remains sensitive even when used with underpowered GWAS datasets. Finally, using ERASE in combination with GWAS data for Parkinsons disease and data on the splicing potential of individual SNPs, we provide evidence to suggest that risk loci for Parkinsons disease are enriched amongst ASEs likely to affect splicing. Thus, we show that ERASE is an important new tool for the integration of ASE and GWAS data, capable of providing novel insights into the pathophysiology of complex diseases.

bioinformatics

Loss-of-function genomic variants with impact on liver-related blood traits highlight potential therapeutic targets for cardiovascular disease

Cardiovascular diseases (CVD), and in particular cerebrovascular and ischemic heart diseases, are leading causes of death globally.1 Lowering circulating lipids is an important treatment strategy to reduce risk.2,3 However, some pharmaceutical mechanisms of reducing CVD may increase risk of fatty liver disease or other metabolic disorders.4,5,6 To identify potential novel therapeutic targets, which may reduce risk of CVD without increasing risk of metabolic disease, we focused on the simultaneous evaluation of quantitative traits related to liver function and CVD. Using a combination of low-coverage (5x) whole-genome sequencing and targeted genotyping, deep genotype imputation based on the TOPMed reference panel7, and genome-wide association study (GWAS) meta-analysis, we analyzed 12 liver-related blood traits (including liver enzymes, blood lipids, and markers of iron metabolism) in up to 203,476 people from three population-based cohorts of different ancestries. We identified 88 likely causal protein-altering variants that were associated with one or more liver-related blood traits. We identified several loss-of-function (LoF) variants reducing low-density lipoprotein cholesterol (LDL-C) or risk of CVD without increased risk of liver disease or diabetes, including variants in known lipid genes (e.g. APOB, LPL). A novel LoF variant, ZNF529:p.K405X, was associated with decreased levels of LDL-C (P=1.3x10-8) but demonstrated no association with liver enzymes or non-fasting blood glucose levels. Silencing of ZNF529 in human hepatocytes resulted in upregulation of LDL receptor (LDLR) and increased LDL-C uptake in the cells, suggesting that inhibition of ZNF529 or its gene product could be used for treating hypercholesterolemia and hence reduce the risk of CVD. Taken together, we demonstrate that simultaneous consideration of multiple phenotypes and a focus on rare protein-altering variants may identify promising therapeutic targets.

genomics

Meta-MultiSKAT: Multiple phenotype meta-analysis for region-based association test

The power of genetic association analyses can be increased by jointly meta-analyzing multiple correlated phenotypes. Here, we develop a meta-analysis framework, Meta-MultiSKAT, that uses summary statistics to test for association between multiple continuous phenotypes and variants in a region of interest. Our approach models the heterogeneity of effects between studies through a kernel matrix and performs a variance component test for association. Using a genotype kernel, our approach can test for rare-variants and the combined effects of both common and rare-variants. To achieve robust power, within Meta-MultiSKAT, we developed fast and accurate omnibus tests combining different models of genetic effects, functional genomic annotations, multiple correlated phenotypes and heterogeneity across studies. Additionally, Meta-MultiSKAT accommodates situations where studies do not share exactly the same set of phenotypes or have differing correlation patterns among the phenotypes. Simulation studies confirm that Meta-MultiSKAT can maintain type-I error rate at exome-wide level of 2.5x10-6. Further simulations under different models of association show that Meta-MultiSKAT can improve power of detection from 23% to 38% on average over single phenotype-based meta-analysis approaches. We demonstrate the utility and improved power of Meta-MultiSKAT in the meta-analyses of four white blood cell subtype traits from the Michigan Genomics Initiative (MGI) and SardiNIA studies.

genetics

Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts

With very large sample sizes, population-based cohorts and biobanks provide an exciting opportunity to identify genetic components of complex traits. To analyze rare variants, gene or region-based multiple variant aggregate tests are commonly used to increase association test power. However, due to the substantial computation cost, existing region-based rare variant tests cannot analyze hundreds of thousands of samples while accounting for confounders, such as population stratification and sample relatedness. Here we propose a scalable generalized mixed model region-based association test that can handle large sample sizes and accounts for unbalanced case-control ratios for binary traits. This method, SAIGE-GENE, utilizes state-of-the-art optimization strategies to reduce computational and memory cost, and hence is applicable to exome-wide and genome-wide region-based analysis for hundreds of thousands of samples. Through the analysis of the HUNT study of 69,716 Norwegian samples and the UK Biobank data of 408,910 White British samples, we show that SAIGE-GENE can efficiently analyze large sample data (N > 400,000) with type I error rates well controlled.

genetics

Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program

Summary paragraphThe Trans-Omics for Precision Medicine (TOPMed) program seeks to elucidate the genetic architecture and disease biology of heart, lung, blood, and sleep disorders, with the ultimate goal of improving diagnosis, treatment, and prevention. The initial phases of the program focus on whole genome sequencing of individuals with rich phenotypic data and diverse backgrounds. Here, we describe TOPMed goals and design as well as resources and early insights from the sequence data. The resources include a variant browser, a genotype imputation panel, and sharing of genomic and phenotypic data via dbGaP. In 53,581 TOPMed samples, >400 million single-nucleotide and insertion/deletion variants were detected by alignment with the reference genome. Additional novel variants are detectable through assembly of unmapped reads and customized analysis in highly variable loci. Among the >400 million variants detected, 97% have frequency <1% and 46% are singletons. These rare variants provide insights into mutational processes and recent human evolutionary history. The nearly complete catalog of genetic variation in TOPMed studies provides unique opportunities for exploring the contributions of rare and non-coding sequence variants to phenotypic variation. Furthermore, combining TOPMed haplotypes with modern imputation methods improves the power and extends the reach of nearly all genome-wide association studies to include variants down to ~0.01% in frequency.

genomics