Search bioRxivSearch

Biology subjects

Daly, M. J.

Publications and source records attributed to Daly, M. J..

At least 19 recordsLinked to original sources

Hidden ‘risk’ in polygenic scores: clinical use today could exacerbate health disparities

Polygenic risk scores (PRS) are poised to improve biomedical outcomes via precision medicine. However, the major ethical and scientific challenge surrounding clinical implementation is that they are many-fold more accurate in European ancestry individuals than others. This disparity is an inescapable consequence of Eurocentric genome-wide association study biases. This highlights that--unlike clinical biomarkers and prescription drugs, which may individually work better in some populations but do not ubiquitously perform far better in European populations--clinical uses of PRS today would systematically afford greater improvement to European descent populations. Early diversifying efforts show promise in levelling this vast imbalance, even when non-European sample sizes are considerably smaller than the largest studies to date. To realize the full and equitable potential of PRS, we must prioritize greater diversity in genetic studies and public dissemination of summary statistics to ensure that health disparities are not increased for those already most underserved.

genetics

Applicability of the mutation-selection balance model to population genetics of heterozygous protein-truncating variants in humans

The fate of alleles in the human population is believed to be highly affected by the stochastic force of genetic drift. Estimation of the strength of natural selection in humans generally necessitates a careful modeling of drift including complex effects of the population history and structure. Protein truncating variants (PTVs) are expected to evolve under strong purifying selection and to have a relatively high per-gene mutation rate. Thus, it is appealing to model the population genetics of PTVs under a simple deterministic mutation-selection balance, as has been proposed earlier [1]. Here, we investigated the limits of this approximation using both computer simulations and data-driven approaches. Our simulations rely on a model of demographic history estimated from 33,370 individual exomes of the Non-Finnish European subset of the ExAC dataset [2]. Additionally, we compared the African and European subset of the ExAC study and analyzed de novo PTVs. We show that the mutation-selection balance model is applicable to the majority of human genes, but not to genes under the weakest selection.

genetics

Contribution of rare and common variants to intellectual disability in a high-risk population sub-isolate of Northern Finland

The contribution of de novo and ultra-rare genetic variants in severe and moderate intellectual disability (ID) has been extensively studied whereas the genetic architecture of mild ID has been less well characterized. To elucidate the genetic background of milder ID we studied a regional cohort of 442 ID patients enriched for mild ID (>50%) from a population isolate of Finland. We analyzed rare variants using exome sequencing and CNV genotyping and common variants using common variant polygenic risk scores. As controls we used a Finnish collection of exome sequenced (n=11311) and GWAS chip genotyped (n=11699) individuals.\n\nWe show that rare damaging variants in genes known to be associated with cognitive defects are observed more often in severe (27%) than in mild ID (13%) patients (p-value: 7.0e-4). We further observed a significant enrichment of protein truncating variants in loss-of-function intolerant genes, as well as damaging missense variants in genes not yet associated with cognitive defects (OR: 2.1, p-value: 3e-8). For the first time to our knowledge, we show that a common variant polygenic load significantly contributes to all severity forms of ID. The heritability explained was the highest for educational attainment (EDU) in mild ID explaining 2.2% of the heritability on liability scale. For more severe ID it was lower at 0.6%. Finally, we identified a homozygote variant in the CRADD gene to be a cause of a specific syndrome with ID and pachygyria. The frequency of this variant is 50x higher in the Finnish population than in non-Finnish Europeans, demonstrating the benefits of utilizing population isolates in rare variant analysis of diseases under negative selection.

genetics

Bayesian model comparison for rare variant association studies of multiple phenotypes

Whole genome sequencing studies applied to large populations or biobanks with extensive phenotyping raise new analytic challenges. The need to consider many variants at a locus or group of genes simultaneously and the potential to study many correlated phenotypes with shared genetic architecture provide opportunities for discovery and inference that are not addressed by the traditional one variant, one phenotype association study. Here, we introduce a Bayesian model comparison approach that we refer to as MRP (Multiple Rare-variants and Phenotypes) for rare-variant association studies that considers correlation, scale, and direction of genetic effects across a group of genetic variants, phenotypes, and studies. The approach requires only summary statistic data. To demonstrate the efficacy of MRP, we apply our method to exome sequencing data (N = 184,698) across 2,019 traits from the UK Biobank, aggregating signals in genes. MRP demonstrates an ability to recover previously-verified signals such as associations between PCSK9 and LDL cholesterol levels. We additionally find MRP effective in conducting meta-analyses in exome data. Notable non-biomarker findings include associations between MC1R and red hair color and skin color, IL17RA and monocyte count, IQGAP2 and mean platelet volume, and JAK2 and platelet count and crit (mass). Finally, we apply MRP in a multi-phenotype setting; after clustering the 35 biomarker phenotypes based on genetic correlation estimates into four clusters, we find that joint analysis of these phenotypes results in substantial power gains for gene-trait associations, such as in TNFRSF13B in one of the clusters containing diabetes and lipid-related traits. Overall, we show that the MRP model comparison approach is able to improve upon useful features from widely-used meta-analysis approaches for rare variant association analyses and prioritize protective modifiers of disease risk.

genetics

A genome-wide association study for shared risk across major psychiatric disorders in a nation-wide birth cohort implicates fetal neurodevelopment as a key mediator

There is mounting evidence that seemingly diverse psychiatric disorders share genetic etiology, but the biological substrates mediating this overlap are not well characterized. Here, we leverage the unique iPSYCH study, a nationally representative cohort ascertained through clinical psychiatric diagnoses indicated in Danish national health registers. We confirm previous reports of individual and cross-disorder SNP-heritability for major psychiatric disorders and perform a cross-disorder genome-wide association study. We identify four novel genome-wide significant loci encompassing variants predicted to regulate genes expressed in radial glia and interneurons in the developing neocortex during midgestation. This epoch is supported by partitioning cross-disorder SNP-heritability which is enriched at regulatory chromatin active during fetal neurodevelopment. These findings indicate that dysregulation of genes that direct neurodevelopment by common genetic variants results in general liability for many later psychiatric outcomes.

genetics

Common variant burden contributes significantly to the familial aggregation of migraine in 1,589 families

It has long been observed that complex traits, including migraine, often aggregate in families, but the underlying genetic architecture behind this is not well understood. Two competing hypotheses exist, emphasizing either rare or common genetic variation. More specifically, familial aggregation could be predominantly explained by rare, penetrant variants that segregate according to Mendelian inheritance or rather by the sufficient polygenic accumulation of many common variants, each with an individually small effect. Some combination of both common and rare variation could also contribute towards a spectrum of disease risk.\n\nWe investigated this in a collection of 8,319 individuals across 1,589 migraine families from Finland. Family members were individually diagnosed by a migraine-specific questionnaire with either migraine without aura (MO, ICHD-3 code 1.1, n=2,357), migraine with typical aura (ICHD- 3 code 1.2.1, n=2,420), hemiplegic migraine (HM, ICHD-3 code 1.2.3, n=540), or no migraine (n=3,002). For comparison, we used population-based migraine cases (n=1,101) and controls (n=13,369) from the FINRISK study. The disease status of FINRISK individuals was assigned based on health registry data from outpatient clinics and/or prescription medication. All individuals were genotyped on the Illumina(R) CoreExome or PsychArray chip platforms and imputed to a Finnish reference panel of 6,962 haplotypes. Polygenic risk scores (PRS), representing the common variant burden in each individual, were calculated using weights from the most recent large-scale genome-wide association study of migraine. To account for family structure in our analyses, we used a mixed-model approach, adjusting for the genetic relationship matrix as a random effect.\n\nWe found a significantly higher common variant burden in familial cases of migraine (for all subtypes, measured by the odds ratio [OR] per standard deviation [SD] increase in PRS; OR = 1.76, 95% CI = 1.71-1.81, P = 1.7x10-109) compared to cases from a population cohort (OR = 1.32, 95% CI = 1.25-1.38, P = 7.2x10-17) when using the population controls as a reference group. The highest enrichment was observed for HM (OR = 1.96, 95% CI = 1.86-2.07, P = 8.7x10-36) and migraine with typical aura (OR = 1.85, 95% CI = 1.79-1.91, P = 1.4x10-86) but enrichment was also present for MO (OR = 1.57, 95% CI = 1.51-1.63, P = 1.1x10-48). Comparing within cases, there was no significant difference in common variant burden between the migraine with aura subtypes, HM and migraine with typical aura (OR = 1.09, 95% CI = 0.99-1.19, P = 0.09), but both showed significantly higher enrichment compared to MO (OR = 1.28, 95% CI = 1.17-1.38, P = 7.3x10-7, and OR = 1.17, 95% CI = 1.11-1.23, P = 4.62x10-5, respectively). Additionally, we found that higher common variant burden corresponded to earlier age of headache onset (OR per SD increase in PRS for 3,631 cases with onset before 20 years old compared to 1,686 cases with onset later than 20 years old; OR = 1.11, 95% CI = 1.05-1.18, P = 8.3x10-4). FINRISK population cases identified from national health registry data were found to have lower common variant burden in comparison to the familial migraine cases (OR = 1.32, 95% CI = 1.25-1.38, P = 6.8x10-17), unless the individuals had attended both a specialist clinic and also received prophylactic migraine treatment (OR = 1.70, 95% CI = 1.53-1.88, P = 3.9x10-9). Finally, although rare variants have been suggested as the primary cause for familial hemiplegic migraine (FHM), we found only four out of 45 sequenced FHM families (8.9%) with a pathogenic mutation in one of the known risk genes.\n\nIn summary, our results demonstrate a substantial contribution of common polygenic variation to familial aggregation in migraine, comparable to both controls and that observed in migraine cases from a population cohort. The findings also suggest that individuals with migraine aura symptoms (either typical aura, which is mostly visual, or rare motor aura) tend to have higher common variant burden on average supporting the polygenic model also in these migraine subtypes.

genetics

Deep coverage whole genome sequences and plasma lipoprotein(a) in individuals of European and African ancestries

Lipoprotein(a), Lp(a), is a modified low-density lipoprotein particle where apolipoprotein(a) (protein product of the LPA gene) is covalently attached to apolipoprotein B. Lp(a) is a highly heritable, causal risk factor for cardiovascular diseases and varies in concentrations across ancestries. To comprehensively delineate the inherited basis for plasma Lp(a), we performed deep-coverage whole genome sequencing in 8,392 individuals of European and African American ancestries. Through whole genome variant discovery and direct genotyping of all structural variants overlapping LPA, we quantified the 5.5kb kringle IV-2 copy number (KIV2-CN), a known LPA structural polymorphism, and developed a model for its imputation. Through common variant analysis, we discovered a novel locus (SORT1) associated with Lp(a)-cholesterol, and also genetic modifiers of KIV2-CN. Furthermore, in contrast to previous GWAS studies, we explain most of the heritability of Lp(a), observing Lp(a) to be 85% heritable among African Americans and 75% among Europeans, yet with notable inter-ethnic heterogeneity. Through analyses of aggregates of rare coding and non-coding variants with Lp(a)-cholesterol, we found the only genome-wide significant signal to be at a non-coding SLC22A3 intronic window also previously described to be associated with Lp(a); however, this association was mitigated by adjustment with KIV2-CN. Finally, using an additional imputation dataset (N=27,344), we performed Mendelian randomization of LPA variant classes, finding that genetically regulated Lp(a) is more strongly associated with incident cardiovascular diseases than directly measured Lp(a), and is significantly associated with measures of subclinical atherosclerosis in African Americans.

genomics

Scaling accurate genetic variant discovery to tens of thousands of samples

Comprehensive disease gene discovery in both common and rare diseases will require the efficient and accurate detection of all classes of genetic variation across tens to hundreds of thousands of human samples. We describe here a novel assembly-based approach to variant calling, the GATK HaplotypeCaller (HC) and Reference Confidence Model (RCM), that determines genotype likelihoods independently per-sample but performs joint calling across all samples within a project simultaneously. We show by calling over 90,000 samples from the Exome Aggregation Consortium (ExAC) that, in contrast to other algorithms, the HC-RCM scales efficiently to very large sample sizes without loss in accuracy; and that the accuracy of indel variant calling is superior in comparison to other algorithms. More importantly, the HC-RCM produces a fully squared-off matrix of genotypes across all samples at every genomic position being investigated. The HC-RCM is a novel, scalable, assembly-based algorithm with abundant applications for population genetics and clinical studies.

genomics

Haplotype sharing provides insights into fine-scale population history and disease in Finland

Finland provides unique opportunities to investigate population and medical genomics because of its adoption of unified national electronic health records, detailed historical and birth records, and serial population bottlenecks. We assemble a comprehensive view of recent population history ([≤]100 generations), the timespan during which most rare disease-causing alleles arose, by comparing pairwise haplotype sharing from 43,254 Finns to geographically and linguistically adjacent countries with different population histories, including 16,060 Swedes, Estonians, Russians, and Hungarians. We find much more extensive sharing in Finns, with at least one [≥] 5 cM tract on average between pairs of unrelated individuals. By coupling haplotype sharing with fine-scale birth records from over 25,000 individuals, we find that while haplotype sharing broadly decays with geographical distance, there are pockets of excess haplotype sharing; individuals from northeast Finland share several-fold more of their genome in identity-by-descent (IBD) segments than individuals from southwest regions containing the major cities of Helsinki and Turku. We estimate recent effective population size changes over time across regions of Finland and find significant differences between the Early and Late Settlement Regions as expected; however, our results indicate more continuous gene flow than previously indicated as Finns migrated towards the northernmost Lapland region. Lastly, we show that haplotype sharing is locally enriched among pairs of individuals sharing rare alleles by an order of magnitude, especially among pairs sharing rare disease causing variants. Our work provides a general framework for using haplotype sharing to reconstruct an integrative view of recent population history and gain insight into the evolutionary origins of rare variants contributing to disease.

genetics

An Unexpectedly Complex Architecture for Skin Pigmentation in Africans

Fewer than 15 genes have been directly associated with skin pigmentation variation in humans, leading to its characterization as a relatively simple trait. However, by assembling a global survey of quantitative skin pigmentation phenotypes, we demonstrate that pigmentation is more complex than previously assumed with genetic architecture varying by latitude. We investigate polygenicity in the Khoe and the San, populations indigenous to southern Africa, who have considerably lighter skin than equatorial Africans. We demonstrate that skin pigmentation is highly heritable, but that known pigmentation loci explain only a small fraction of the variance. Rather, baseline skin pigmentation is a complex, polygenic trait in the KhoeSan. Despite this, we identify canonical and non-canonical skin pigmentation loci, including near SLC24A5, TYRP1, SMARCA2/VLDLR, and SNX13 using a genome-wide association approach complemented by targeted resequencing. By considering diverse, under-studied African populations, we show how the architecture of skin pigmentation can vary across humans subject to different local evolutionary pressures.\n\nHighlightsO_LISkin pigmentation in Africans is far more polygenic than light skin pigmentation in Eurasians.\nC_LIO_LIKhoeSan[§] populations, which diverged early in human prehistory from other populations, have lightened skin pigmentation compared to equatorial Africans.\nC_LIO_LISkin color is highly heritable in the KhoeSan, but pigmentation variability is not well explained by previously discovered pigmentation genes.\nC_LIO_LIWe perform the first GWAS for pigmentation in African KhoeSan populations and identify canonical pigmentation loci near TYRP1 and in SLC24A5, as well as novel associations surrounding SMARCA2 and other genes.\nC_LI

genetics

Paternal-age-related de novo mutations and risk for five disorders

BackgroundThere are well-established epidemiologic associations between advanced paternal age and increased offspring risk for several psychiatric and developmental disorders. These associations are commonly attributed to age-related de novo mutations. However, the actual magnitude of risk conferred by age-related de novo mutations in the male germline is unknown. Quantifying this risk would clarify the clinical and public health significance of delayed paternity.\n\nMethodsUsing results from large, parent-child trio whole-exome-sequencing studies, we estimated the relationship between paternal-age-related de novo single nucleotide variants (dnSNVs) and offspring risk for five disorders: autism spectrum disorders (ASD), congenital heart disease (CHD), neurodevelopmental disorders with epilepsy (EPI), intellectual disability (ID), and schizophrenia (SCZ). Using Danish national registry data, we then investigated the degree to which the epidemiologic association between each disorder and advanced paternal age was consistent with the estimated role of de novo mutations.\n\nResultsIncidence rate ratios comparing dnSNV-based risk to offspring of 45 versus 25-year-old fathers ranged from 1.05 (95% confidence interval 1.01-1.13) for SCZ to 1.29 (95% CI 1.13-1.68) for ID. Epidemiologic estimates of paternal age risk for CHD, ID and EPI were consistent with the dnSNV effect. However, epidemiologic effects for ASDs and SCZ significantly exceeded the risk that could be explained by dnSNVs alone (p<2e-4 for both comparisons).\n\nConclusionIncreasing dnSNVs due to advanced paternal age confer a small amount of offspring risk for psychiatric and developmental disorders. For ASD and SCZ, epidemiologic associations with delayed paternity largely reflect factors that cannot be assumed to increase with age.

genetics

Medical relevance of protein-truncating variants across 337,208 individuals in the UK Biobank study

Protein-truncating variants can have profound effects on gene function and are critical for clinical genome interpretation and generating therapeutic hypotheses, but their relevance to medical phenotypes has not been systematically assessed. We characterized the effect of 18,228 protein-truncating variants across 135 phenotypes from the UK Biobank and found 27 associations between medical phenotypes and protein-truncating variants in genes outside the major histocompatibility complex. We performed phenome-wide analyses and directly measured the effect of homozygous carriers, commonly referred to as \"human knockouts,\" across medical phenotypes for genes implicated to be protective against disease or associated with at least one phenotype in our study and found several genes with strong pleiotropic or non-additive effects. Our results illustrate the importance of protein-truncating variants in a variety of diseases.

genetics

A Strategy for Large-Scale Systematic Pan-Cancer Germline Rare Variation Analysis

Traditionally, genetic studies in cancer are focused on somatic mutations found in tumors and absent from the normal tissue. Identification of shared attributes in germline variation could aid discrimination of high-risk from likely benign mutations and narrow the search space for new cancer predisposing genes. Extraordinary progress made in analysis of common variation with GWAS methodology does not provide sufficient resolution to understand rare variation. To fulfil missing classification for rare germline variation we assembled datasets of whole exome sequences from >2,000 patients with different types of cancers: breast cancer, colon cancer and cutaneous and ocular melanomas matched to more than 7,000 non-cancer controls and analyzed germline variation in known cancer predisposing genes to identify common properties of disease associated mutations and new candidate cancer susceptibility genes. Lists of all cancer predisposing genes were divided into subclasses according to the mode of inheritance of the related cancer syndrome or contribution to known major cancer pathways. Out of all subclasses only genes linked to dominant syndromes presented significant rare germline variants enrichment in cases. Separate analysis of protein-truncating and missense variation in this subclass of genes confirmed significant prevalence of protein-truncating variants in cases only in loss-of-function tolerant genes (pLI<0.1), while ultra-rare missense mutations were significantly overrepresented in cases only in constrained genes (pLI>0.9). Taken together, our findings provide insights into the distribution and types of mutations underlying inherited cancer predisposition.

genetics

Gene family information facilitates variant interpretation and identification of disease-associated genes

Differentiating risk-conferring from benign missense variants, and therefore optimal calculation of gene-variant burden, represent a major challenge in particular for rare and genetic heterogeneous disorders. While orthologous gene conservation is commonly employed in variant annotation, approximately 80% of known disease-associated genes are paralogs and belong to gene families. It has not been thoroughly investigated how gene family information can be utilized for disease gene discovery and variant interpretation. We developed a paralog conservation score to empirically evaluate whether paralog conserved or nonconserved sites of in-human paralogs are important for protein function. Using this score, we demonstrate that disease-associated missense variants are significantly enriched at paralog conserved sites across all disease groups and disease inheritance models tested. Next, we assessed whether gene family information could assist in discovering novel disease-associated genes. We subsequently developed a gene family de novo enrichment framework that identified 43 exome-wide enriched gene families including 98 de novo variant carrying genes in more than 10k neurodevelopmental disorder patients. 33 gene family enriched genes represent novel candidate genes which are brain expressed and variant constrained in neurodevelopmental disorders.

genetics

Regional missense constraint improves variant deleteriousness prediction

Given increasing numbers of patients who are undergoing exome or genome sequencing, it is critical to establish tools and methods to interpret the impact of genetic variation. While the ability to predict deleteriousness for any given variant is limited, missense variants remain a particularly challenging class of variation to interpret, since they can have drastically different effects depending on both the precise location and specific amino acid substitution of the variant. In order to better evaluate missense variation, we leveraged the exome sequencing data of 60,706 individuals from the Exome Aggregation Consortium (ExAC) dataset to identify sub-genic regions that are depleted of missense variation. We further used this depletion as part of a novel missense deleteriousness metric named MPC. We applied MPC to de novo missense variants and identified a category of de novo missense variants with the same impact on neurodevelopmental disorders as truncating mutations in intolerant genes, supporting the value of incorporating regional missense constraint in variant interpretation.

genomics

Base-Specific Mutational Intolerance Near Splice-Sites Clarifies Role Of Non-Essential Splice Nucleotides

Variation in RNA splicing (i.e., alternative splicing) plays an important role in many diseases. Variants near 5' and 3' splice sites often affect splicing, but the effects of these variants on splicing and disease have not been fully characterized beyond the 2 \"essential\" splice nucleotides flanking each exon. Here we provide quantitative measurements of tolerance to mutational disruptions by position and reference allele-alternative allele combination. We show that certain reference alleles are particularly sensitive to mutations, regardless of the alternative alleles into which they are mutated. Using public RNA-seq data, we demonstrate that individuals carrying such variants have significantly lower levels of the correctly spliced transcript compared to individuals without them, and confirm that these specific substitutions are highly enriched for known Mendelian mutations. Our results propose a more refined definition of the \"splice region\" and offer a new way to prioritize and provide functional interpretation of variants identified in diagnostic sequencing and association studies.

genomics

Limited contribution of rare, noncoding variation to autism spectrum disorder from sequencing of 2,076 genomes in quartet families

Genomic studies to date in autism spectrum disorder (ASD) have largely focused on newly arising mutations that disrupt protein coding sequence and strongly influence risk. We evaluate the contribution of noncoding regulatory variation across the size and frequency spectrum through whole genome sequencing of 519 ASD cases, their unaffected sibling controls, and parents. Cases carry a small excess of de novo (1.02-fold) noncoding variants, which is not significant after correcting for paternal age. Assessing 51,801 regulatory classes, no category is significantly associated with ASD after correction for multiple testing. The strongest signals are observed in coding regions, including structural variation not detected by previous technologies and missense variation. While rare noncoding variation likely contributes to risk in neurodevelopmental disorders, no category of variation has impact equivalent to loss-of-function mutations. Average effect sizes are likely to be smaller than that for coding variation, requiring substantially larger samples to quantify this risk.

genomics