Search bioRxivSearch

Biology subjects

Ganna, A.

Publications and source records attributed to Ganna, A..

6 recordsLinked to original sources

Signals of polygenic adaptation on height have been overestimated due to uncorrected population structure in genome-wide association studies

Genetic predictions of height differ among human populations and these differences are too large to be explained by genetic drift. This observation has been interpreted as evidence of polygenic adaptation. Differences across populations were detected using SNPs genome-wide significantly associated with height, and many studies also found that the signals grew stronger when large numbers of subsignificant SNPs were analyzed. This has led to excitement about the prospect of analyzing large fractions of the genome to detect subtle signals of selection and claims of polygenic adaptation for multiple traits. Polygenic adaptation studies of height have been based on SNP effect size measurements in the GIANT Consortium meta-analysis. Here we repeat the height analyses in the UK Biobank, a much more homogeneously designed study. Our results show that polygenic adaptation signals based on large numbers of SNPs below genome-wide significance are extremely sensitive to biases due to uncorrected population structure.

evolutionary biology

Low-frequency variant functional architectures reveal strength of negative selection across coding and non-coding annotations

Common variant heritability is known to be concentrated in variants within cell-type-specific non-coding functional annotations, with a limited role for common coding variants. However, little is known about the functional distribution of low-frequency variant heritability. Here, we partitioned the heritability of both low-frequency (0.5% [&le;] MAF < 5%) and common (MAF [&ge;] 5%) variants in 40 UK Biobank traits (average N = 363K) across a broad set of coding and non-coding functional annotations, employing an extension of stratified LD score regression to low-frequency variants that produces robust results in simulations. We determined that non-synonymous coding variants explain 17{+/-}1% of low-frequency variant heritability [Formula] versus only 2.1{+/-}0.2% of common variant heritability [Formula], and that regions conserved in primates explain nearly half of [Formula] (43{+/-}2%). Other annotations previously linked to negative selection, including non-synonymous variants with high PolyPhen-2 scores, non-synonymous variants in genes under strong selection, and low-LD variants, were also significantly more enriched for [Formula] as compared to [Formula]. Cell-type-specific non-coding annotations that were significantly enriched for [Formula] of corresponding traits tended to be similarly enriched for [Formula] for most traits, but more enriched for brain-related annotations and traits. For example, H3K4me3 marks in brain DPFC explain 57{+/-}12% of [Formula] vs. 12{+/-}2% of [Formula] for neuroticism, implicating the action of negative selection on low-frequency variants affecting gene regulation in the brain. Forward simulations confirmed that the ratio of low-frequency variant enrichment vs. common variant enrichment primarily depends on the mean selection coefficient of causal variants in the annotation, and can be used to predict the effect size variance of causal rare variants (MAF < 0.5%) in the annotation, informing their prioritization in whole-genome sequencing studies. Our results provide a deeper understanding of low-frequency variant functional architectures and guidelines for the design of association studies targeting functional classes of low-frequency and rare variants.

genetics

Deep coverage whole genome sequences and plasma lipoprotein(a) in individuals of European and African ancestries

Lipoprotein(a), Lp(a), is a modified low-density lipoprotein particle where apolipoprotein(a) (protein product of the LPA gene) is covalently attached to apolipoprotein B. Lp(a) is a highly heritable, causal risk factor for cardiovascular diseases and varies in concentrations across ancestries. To comprehensively delineate the inherited basis for plasma Lp(a), we performed deep-coverage whole genome sequencing in 8,392 individuals of European and African American ancestries. Through whole genome variant discovery and direct genotyping of all structural variants overlapping LPA, we quantified the 5.5kb kringle IV-2 copy number (KIV2-CN), a known LPA structural polymorphism, and developed a model for its imputation. Through common variant analysis, we discovered a novel locus (SORT1) associated with Lp(a)-cholesterol, and also genetic modifiers of KIV2-CN. Furthermore, in contrast to previous GWAS studies, we explain most of the heritability of Lp(a), observing Lp(a) to be 85% heritable among African Americans and 75% among Europeans, yet with notable inter-ethnic heterogeneity. Through analyses of aggregates of rare coding and non-coding variants with Lp(a)-cholesterol, we found the only genome-wide significant signal to be at a non-coding SLC22A3 intronic window also previously described to be associated with Lp(a); however, this association was mitigated by adjustment with KIV2-CN. Finally, using an additional imputation dataset (N=27,344), we performed Mendelian randomization of LPA variant classes, finding that genetically regulated Lp(a) is more strongly associated with incident cardiovascular diseases than directly measured Lp(a), and is significantly associated with measures of subclinical atherosclerosis in African Americans.

genomics

Deep-coverage whole genome sequences and blood lipids among 16,324 individuals

Deep-coverage whole genome sequencing at the population level is now feasible and offers potential advantages for locus discovery, particularly in the analysis rare mutations in non-coding regions. Here, we performed whole genome sequencing in 16,324 participants from four ancestries at mean depth >29X and analyzed correlations of genotypes with four quantitative traits - plasma levels of total cholesterol, low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol, and triglycerides. We conducted a discovery analysis including common or rare variants in coding as well as non-coding regions and developed a framework to interpret genome sequence for dyslipidemia risk. Common variant association yielded loci previously described with the exception of a few variants not captured earlier by arrays or imputation. In coding sequence, rare variant association yielded known Mendelian dyslipidemia genes and, in non-coding sequence, we detected no rare variant association signals after application of four approaches to aggregate variants in non-coding regions. We developed a new, genome-wide polygenic score for LDL-C and observed that a high polygenic score conferred similar effect size to a monogenic mutation (~30 mg/dl higher LDL-C for each); however, among those with extremely high LDL-C, a high polygenic score was considerably more prevalent than a monogenic mutation (23% versus 2% of participants, respectively).

genomics

Quantifying the impact of rare and ultra-rare coding variation across the phenotypic spectrum

There is a limited understanding about the impact of rare protein truncating variants across multiple phenotypes. We explore the impact of this class of variants on 13 quantitative traits and 10 diseases using whole-exome sequencing data from 100,296 individuals. Protein truncating variants in genes intolerant to this class of mutations increased risk of autism, schizophrenia, bipolar disorder, intellectual disability, ADHD. In individuals without these disorders, there was an association with shorter height, lower education, increased hospitalization and reduced age. Gene sets implicated from GWAS did not show a significant protein truncating variants-burden beyond what captured by established Mendelian genes. In conclusion, we provide the most thorough investigation to date of the impact of rare deleterious coding variants on complex traits, suggesting widespread pleiotropic risk.\n\nMain abbreviations

genetics

Heterogeneous Contribution of Microdeletions in the Development of Common Generalized and Focal epilepsies

BackgroundMicrodeletions are known to confer risk to epilepsy, particularly at genomic rearrangement \"hotspot\" loci. However, deciphering their role outside hotspots and risk assessment by epilepsy sub-type has not been conducted.\n\nMethodsWe assessed the burden, frequency and genomic content of rare, large microdeletions found in a previously published cohort of 1,366 patients with Genetic Generalized Epilepsy (GGE) plus two sets of additional unpublished genome-wide microdeletions found in 281 Rolandic Epilepsy (RE) and 807 Adult Focal Epilepsy (AFE) patients, totaling 2,454 cases. These microdeletion sets were assessed in a combined analysis and in sub-type specific approaches against 6,746 ethnically matched controls.\n\nResultsWhen hotspots are considered, we detected an enrichment of microdeletions in the combined epilepsy analysis (adjusted-P= 2.00x10-7; OR = 1.89; 95%-CI: 1.51-2.35), where the implicated microdeletions overlapped with rarely deleted genes and those involved in neurodevelopmental processes. Sub-type specific analyses showed that hotspot deletions in the GGE subgroup contribute most of the signal (adjusted-P = 1.22x10-12; OR = 7.45; 95%-CI = 4.20-11.97). Outside hotspot loci, microdeletions were enriched in the GGE cohort for neurodevelopmental genes (adjusted-P = 4.78x10-3; OR = 2.30; 95%-CI = 1.42-3.70), whereas no additional signal was observed for RE and AFE. Still, gene content analysis was able to identify known (NRXN1, RBFOX1 and PCDH7) and novel (LOC102723362) candidate genes affected in more than one epilepsy sub-type but not in controls.\n\nConclusionsOur results show a heterogeneous effect of recurrent and non-recurrent microdeletions as part of the genetic architecture of GGE and a minor to negligible contribution in the etiology of RE and AFE.

genetics