Search bioRxivSearch

Biology subjects

Mark J Daly

Publications and source records attributed to Mark J Daly.

13 recordsLinked to original sources

Population genetic history and polygenic risk biases in 1000 Genomes populations

The vast majority of genome-wide association studies are performed in Europeans, and their transferability to other populations is dependent on many factors (e.g. linkage disequilibrium, allele frequencies, genetic architecture). As medical genomics studies become increasingly large and diverse, gaining insights into population history and consequently the transferability of disease risk measurement is critical. Here, we disentangle recent population history in the widely-used 1000 Genomes Project reference panel, with an emphasis on populations underrepresented in medical studies. To examine the transferability of single-ancestry GWAS, we used published summary statistics to calculate polygenic risk scores for six well-studied traits and diseases. We identified directional inconsistencies in all scores; for example, height is predicted to decrease with genetic distance from Europeans, despite robust anthropological evidence that West Africans are as tall as Europeans on average. To gain deeper quantitative insights into GWAS transferability, we developed a complex trait coalescent-based simulation framework considering effects of polygenicity, causal allele frequency divergence, and heritability. As expected, correlations between true and inferred risk were typically highest in the population from which summary statistics were derived. We demonstrated that scores inferred from European GWAS were biased by genetic drift in other populations even when choosing the same causal variants, and that biases in any direction were possible and unpredictable. This work cautions that summarizing findings from large-scale GWAS may have limited portability to other populations using standard approaches, and highlights the need for generalized risk prediction methods and the inclusion of more diverse individuals in medical genomics.

Genomics

The ExAC Browser: Displaying reference data information from over 60,000 exomes

Worldwide, hundreds of thousands of humans have had their genomes or exomes sequenced, and access to the resulting data sets can provide valuable information for variant interpretation and understanding gene function. Here, we present a lightweight, flexible browser framework to display large population datasets of genetic variation. We demonstrate its use for exome sequence data from 60,706 individuals in the Exome Aggregation Consortium (ExAC). The ExAC browser provides gene- and transcript-centric displays of variation, a critical view for clinical applications. Additionally, we provide a variant display, which includes population frequency and functional annotation data as well as short read support for the called variant. This browser is open-source, freely available, and has already been used extensively by clinical laboratories worldwide.

Genomics

Bootstrat: Population Informed Bootstrapping for Rare Variant Tests

Recent advances in genotyping and sequencing technologies have made detecting rare variants in large cohorts possible. Various analytic methods for associating disease to rare variants have been proposed, including burden tests, C-alpha and SKAT. Most of these methods, however, assume that samples come from a homogeneous population, which is not realistic for analyses of large samples. Not correcting for population stratification causes inflated p-values and false-positive associations. Here we propose a population-informed bootstrap resampling method that controls for population stratification (Bootstrat) in rare variant tests. In essence, the Bootstrat procedure uses genetic distance to create a phenotype probability for each sample. We show that this empirical approach can effectively correct for population stratification while maintaining statistical power comparable to established methods of controlling for population stratification. The Bootstrat scheme can be easily applied to existing rare variant testing methods with reasonable computational complexity.\n\nAuthor SummaryRecent technology advances have enabled large-scale analysis of rare variants, but properly testing rare variants remains a significant challenge as most rare variant testing methods assume a sample of homogenous ethnicity, an assumption often not true for large cohorts. Failure to account for this heterogeneity increases the type I error rate. Here we propose a bootstrap scheme applicable to most existing rare variant testing methods to control for population heterogeneity. This scheme uses a randomization layer to establish a null distribution of the test statistics while preserving the sample genetic relationships. The null distribution is then used to calculate an empirical p-value that accounts for population heterogeneity. We demonstrate how this scheme successfully controls the type I error rate without loss of statistical power.

Genomics

Ultra-rare disruptive and damaging mutations influence educational attainment in the general population

Ultra-rare inherited and de novo disruptive variants in highly constrained (HC) genes are enriched in neurodevelopmental disorders 1-5. However, their impact on cognition in the general population has not been explored. We hypothesize that disruptive and damaging ultra-rare variants (URVs) in HC genes not only confer risk to neurodevelopmental disorders, but also influence general cognitive abilities measured indirectly by years of education (YOE). We tested this hypothesis in 14,133 individuals with whole exome or genome sequencing data. The presence of one or more URVs was associated with a decrease in YOE (3.1 months less for each additional mutation; P-value=3.3x10-8) and the effect was stronger in HC genes enriched for brain expression (6.5 months less, P-value=3.4x10-5). The effect of these variants was more pronounced than the estimated effects of runs of homozygosity and pathogenic copy number variation 6-9. Our findings suggest that effects of URVs in HC genes are not confined to severe neurodevelopmental disorder, but influence the cognitive spectrum in the general population

Genetics

A method to exploit the structure of genetic ancestry space to enhance case-control studies

One goal of human genetics is to understand the genetic basis of disease, a challenge for diseases of complex inheritance because risk alleles are few relative to the vast set of benign variants. Risk variants are often sought by association studies in which allele frequencies in cases are contrasted with those from population-based samples used as controls. In an ideal world we would know population-level allele frequencies, releasing researchers to focus on case subjects. We argue this ideal is possible, at least theoretically, and we outline a path to achieving it in reality. If such a resource were to exist, it would yield ample savings and would facilitate the effective use of data repositories by removing administrative and technical barriers. We call this concept the Universal Control Repository Network (UNICORN), a means to perform association analyses without necessitating direct access to individual-level control data. Our approach to UNICORN uses existing genetic resources and various statistical tools to analyze these data, including hierarchical clustering with spectral analysis of ancestry; and empirical Bayesian analysis along with Gaussian spatial processes to estimate ancestry-specific allele frequencies. We demonstrate our approach using tens of thousands of controls from studies of Crohns disease, showing how it controls false positives, provides power similar to that achieved when all control data are directly accessible, and enhances power when control data are limiting or even imperfectly matched ancestrally. These results highlight how UNICORN can enable reliable, powerful and convenient genetic association analyses without access to the individual level data.

Genomics

Polygenic risk for schizophrenia is associated with social cognition across development

Breakthroughs in genomics have begun to unravel the genetic architecture of schizophrenia risk, providing methods for quantifying schizophrenia polygenic risk based on common genetic variants. Our objective in the current study was to understand the relationship between schizophrenia genetic risk variants and neurocognitive development in healthy individuals. We first used combined genomic and neurocognitive data from the Philadelphia Neurodevelopmental Cohort (PNC; 4303 participants ages 8 - 21 years) to screen 26 neurocognitive phenotypes for their association with schizophrenia polygenic risk.\n\nSchizophrenia polygenic risk was estimated for each participant based on summary statistics from the most recent schizophrenia genome-wide association analysis (Psychiatric Genomics Consortium 2014). After correction for multiple comparisons, greater schizophrenia polygenic risk was significantly associated with reduced speed of emotion identification and verbal reasoning. These associations were significant by age 9 and there was no evidence of interaction between schizophrenia polygenic risk and age on neurocognitive performance. We then looked at the association between schizophrenia polygenic risk and emotion identification speed in the Harvard / MGH Brain Genomics Superstruct Project sample (GSP; 695 participants age 18-35 years), where we replicated the association between schizophrenia polygenic risk and emotion identification speed. These analyses provide evidence for a replicable association between polygenic risk for schizophrenia and specific aspects of neurocognitive performance. Our findings indicate that individual differences in genetic risk for schizophrenia are linked with the development of social cognition and potentially verbal reasoning, and that these associations emerge relatively early in development.

Genetics

Novel protective associations with age-related macular degeneration: A common variant near CTRB1 and a rare variant in PELI3

Although >20 common frequency age-related macular degeneration (AMD) alleles have been discovered with genome-wide association studies, substantial disease heritability remains unexplained. In this study we sought to identify additional variants, both common and rare, that have an association with advanced AMD. We genotyped 4,332 cases and 25,268 controls of European ancestry from three different populations using the Illumina Infinium HumanExome BeadChip. We performed meta-analyses to identify associations with common variants and performed single variant and gene-based burden tests to identify associations with rare variants. Two protective, low frequency, non-synonymous variants A307V in PELI3 (odds ratio [OR]=0.14, P=4.3x10-10) and N1050Y in CFH (OR=0.76, Pconditional=1.6x10-11) were significantly associated with a decrease in risk of AMD. Additionally, we identified an enrichment of protective alleles in PELI3 using a burden test (OR=0.14). The new variants have a large effect size, similar to rare mutations we reported previously in a targeted sequencing study, which remain significant in this analysis: CFH R1210C (OR=18.82, P=3.5x10-07), C3 K155Q (OR=3.27, P=1.5x10-10), and C9 P167S (OR=2.04, P=2.8x10-07). We also identified a strong protective signal for a common variant (rs8056814) near CTRB1 associated with a decrease in AMD risk (logistic regression: OR = 0.71, P = 1.8x10-07; Firth corrected OR = 0.64, P = 9.6x10-11). This study supports the involvement of both common and low frequency protective variants in AMD. It also may expand the role of the high-density lipoprotein pathway and branches of the innate immune pathway, outside that of the complement system, in the etiology of AMD.

Genetics

Human knockouts in a cohort with a high rate of consanguinity

A major goal of biomedicine is to understand the function of every gene in the human genome.1 Null mutations can disrupt both copies of a given gene in humans and phenotypic analysis of such human knockouts can provide insight into gene function. To date, comprehensive analysis of genes knocked out in humans has been limited by the fact that null mutations are infrequent in the general population and so, observing an individual homozygous null for a given gene is exceedingly rare.2,3 However, consanguineous unions are more likely to result in offspring who carry homozygous null mutations. In Pakistan, consanguinity rates are notably high.4 Here, we sequenced the protein-coding regions of 7,078 adult participants living in Pakistan and performed phenotypic analysis to identify homozygous null individuals and to understand consequences of complete gene disruption in humans. We enumerated 36,850 rare (<1 % minor allele frequency) null mutations. These homozygous null mutations led to complete inactivation of 961 genes in at least one participant. Homozygosity for null mutations at APOC3 was associated with absent plasma apolipoprotein C-III levels; at PLAG27, with absent enzymatic activity of soluble lipoprotein-associated phospholipase A2; at CYP2F1, with higher plasma interleukin-8 concentrations; and at either A3GALT2 or NRG4, with markedly reduced plasma insulin C-peptide concentrations. After physiologic challenge with oral fat, APOC3 knockouts displayed marked blunting of the usual post-prandial rise in plasma triglycerides compared to wild-type family members. These observations provide a roadmap to understand the consequences of complete disruption of a large fraction of genes in the human genome.

Genomics

Genetic risk for autism spectrum disorders and neuropsychiatric variation in the general population

Almost all genetic risk factors for autism spectrum disorders (ASDs) can be found in the general population, but the effects of that risk are unclear in people not ascertained for neuropsychiatric symptoms. Using several large ASD consortia and population based resources, we find genetic links between ASDs and typical variation in social behavior and adaptive functioning. This finding is evidenced through both inherited and de novo variation, indicating that multiple types of genetic risk for ASDs influence a continuum of behavioral and developmental traits, the severe tail of which can result in an ASD or other neuropsychiatric disorder diagnosis. A continuum model should inform the design and interpretation of studies of neuropsychiatric disease biology.

Genetics

Abundant contribution of short tandem repeats to gene expression variation in humans

Expression quantitative trait loci (eQTLs) are a key tool to dissect cellular processes mediating complex diseases. However, little is known about the role of repetitive elements as eQTLs. We report a genome-wide survey of the contribution of Short Tandem Repeats (STRs), one of the most polymorphic and abundant repeat classes, to gene expression in humans. Our survey identified 2,060 significant expression STRs (eSTRs). These eSTRs were replicable in orthogonal populations and expression assays. We used variance partitioning to disentangle the contribution of eSTRs from linked SNPs and indels and found that eSTRs contribute 10%-15% of the cis-heritability mediated by all common variants. Functional genomic analyses showed that eSTRs are enriched in conserved regions, co-localize with regulatory elements, and are predicted to modulate histone modifications. Our results show that eSTRs provide a novel set of regulatory variants and highlight the contribution of repeats to the genetic architecture of quantitative human traits.

Genomics

Network analysis of genome-wide selective constraint reveals a gene network active in early fetal brain intolerant of mutation

Using robust, integrated analysis of multiple genomic datasets, we show that genes depleted for non-synonymous de novo mutations form a subnetwork of 72 members under strong selective constraint. We further show this subnetwork is preferentially expressed in the early development of the human hippocampus and is enriched for genes mutated in neurological, but not other, Mendelian disorders. We thus conclude that carefully orchestrated developmental processes are under strong constraint in early brain development, and perturbations caused by mutation have adverse outcomes subject to strong purifying selection. Our findings demonstrate that selective forces can act on groups of genes involved in the same process, supporting the notion that adaptation can act coordinately on multiple genes. Our approach provides a statistically robust, interpretable way to identify the tissues and developmental times where groups of disease genes are active. Our findings highlight the importance of considering the interactions between genes when analyzing genome-wide sequence data.

Genetics

An Atlas of Genetic Correlations across Human Diseases and Traits

Identifying genetic correlations between complex traits and diseases can provide useful etiological insights and help prioritize likely causal relationships. The major challenges preventing estimation of genetic correlation from genome-wide association study (GWAS) data with current methods are the lack of availability of individual genotype data and widespread sample overlap among meta-analyses. We circumvent these difficulties by introducing a technique for estimating genetic correlation that requires only GWAS summary statistics and is not biased by sample overlap. We use our method to estimate 300 genetic correlations among 25 traits, totaling more than 1.5 million unique phenotype measurements. Our results include genetic correlations between anorexia nervosa and schizophrenia, anorexia and obesity and associations between educational attainment and several diseases. These results highlight the power of genome-wide analyses, since there currently are no genome-wide significant SNPs for anorexia nervosa and only three for educational attainment.

Genomics

LD Score Regression Distinguishes Confounding from Polygenicity in Genome-Wide Association Studies

Both polygenicity1,2 (i.e. many small genetic effects) and confounding biases, such as cryptic relatedness and population stratification3, can yield inflated distributions of test statistics in genome-wide association studies (GWAS). However, current methods cannot distinguish between inflation from bias and true signal from polygenicity. We have developed an approach that quantifies the contributions of each by examining the relationship between test statistics and linkage disequilibrium (LD). We term this approach LD Score regression. LD Score regression provides an upper bound on the contribution of confounding bias to the observed inflation in test statistics and can be used to estimate a more powerful correction factor than genomic control4-14. We find strong evidence that polygenicity accounts for the majority of test statistic inflation in many GWAS of large sample size.

Genomics