Search bioRxivSearch

Biology subjects

Price, A.

Publications and source records attributed to Price, A..

11 recordsLinked to original sources

Rheumatoid arthritis heritability is concentrated in regulatory elements with CD4+ T cell-state-specific transcription factor binding profiles

Despite significant progress in annotating the genome with experimental methods, much of the regulatory noncoding genome remains poorly defined. Here we assert that regulatory elements may be characterized by leveraging local epigenomic signatures at sites where specific transcription factors (TFs) are bound. To link these two identifying features, we introduce IMPACT, a genome annotation strategy which identifies regulatory elements defined by cell-state-specific TF binding profiles, learned from 515 chromatin and sequence annotations. We validate IMPACT using multiple compelling applications. First, IMPACT predicts TF motif binding with high accuracy (average AUC 0.92, s.e. 0.03; across 8 TFs), a significant improvement (all p<6.9e-15) over intersecting motifs with open chromatin (average AUC 0.66, s.e. 0.11). Second, an IMPACT annotation trained on RNA polymerase II is more enriched for peripheral blood cis-eQTL variation (N=3,754) than sequence based annotations, such as promoters and regions around the TSS, (permutation p<1e-3, 25% average increase in enrichment). Third, integration with rheumatoid arthritis (RA) summary statistics from European (N=38,242) and East Asian (N=22,515) populations revealed that the top 5% of CD4+ Treg IMPACT regulatory elements capture 85.7% (s.e. 19.4%) of RA h2 (p<1.6e-5) and that the top 9.8% of Treg IMPACT regulatory elements, consisting of all SNPs with a non-zero annotation value, capture 97.3% (s.e. 18.2%) of RA h2 (p<7.6e-7), the most comprehensive explanation for RA h2 to date. In comparison, the average RA h2 captured by compared CD4+ T histone marks is 42.3% and by CD4+ T specifically expressed gene sets is 36.4%. Finally, integration with RA fine-mapping data (N=27,345) revealed a significant enrichment (2.87, p<8.6e-3) of putatively causal variants across 20 RA associated loci in the top 1% of CD4+ Treg IMPACT regulatory regions. Overall, we find that IMPACT generalizes well to other cell types in identifying complex trait associated regulatory elements.

genetics

CAN NEUROPATHIC PAIN PREDICT RESPONSE TO ARTHROPLASTY IN KNEE OSTEOARTHRITIS? A PROSPECTIVE OBSERVATIONAL COHORT STUDY.

A significant proportion of patients with knee osteoarthritis (OA) continue to have severe ongoing pain following knee replacement surgery. Central sensitization and features suggestive of neuropathic pain before surgery may result in a poor outcome post-operatively. In this prospective observational study of patients undergoing primary knee arthroplasty (n=120), the modified PainDETECT score was used to divide patients, with primary knee OA, into nociceptive (<13), unclear (13-18) and neuropathic -like pain (>18) groups pre-operatively. Response to surgery was compared between groups using the Oxford Knee Score (OKS) and the presence of moderate to severe long-term pain 12 months after arthroplasty. The analyses were replicated in a larger independent cohort study (n=404). 120 patients were recruited to the main study cohort: 63 (52%) nociceptive pain; 32 (27%) unclear pain; 25 (21%) neuropathic-like pain. Patients with neuropathic-like pain had significantly worse OKS pre and post-operatively, compared to the nociceptive pain group, independent of age, sex and BMI. At 12-months post-operatively the mean OKS was 4 points lower in the neuropathic-like group compared with the nociceptive group in the study cohort (non-significant); with a difference of 5 points in the replication cohort (p<0.001). Moderate to severe long-term pain after arthroplasty at 12-months was present in 50% of the neuropathic-like pain group versus 24% in the nociceptive pain group, in the replication cohort (p<0.001). Neuropathic pain is common and targeted therapy pre, peri and post-operatively may improve treatment response.

epidemiology

Low-frequency variant functional architectures reveal strength of negative selection across coding and non-coding annotations

Common variant heritability is known to be concentrated in variants within cell-type-specific non-coding functional annotations, with a limited role for common coding variants. However, little is known about the functional distribution of low-frequency variant heritability. Here, we partitioned the heritability of both low-frequency (0.5% [&le;] MAF < 5%) and common (MAF [&ge;] 5%) variants in 40 UK Biobank traits (average N = 363K) across a broad set of coding and non-coding functional annotations, employing an extension of stratified LD score regression to low-frequency variants that produces robust results in simulations. We determined that non-synonymous coding variants explain 17{+/-}1% of low-frequency variant heritability [Formula] versus only 2.1{+/-}0.2% of common variant heritability [Formula], and that regions conserved in primates explain nearly half of [Formula] (43{+/-}2%). Other annotations previously linked to negative selection, including non-synonymous variants with high PolyPhen-2 scores, non-synonymous variants in genes under strong selection, and low-LD variants, were also significantly more enriched for [Formula] as compared to [Formula]. Cell-type-specific non-coding annotations that were significantly enriched for [Formula] of corresponding traits tended to be similarly enriched for [Formula] for most traits, but more enriched for brain-related annotations and traits. For example, H3K4me3 marks in brain DPFC explain 57{+/-}12% of [Formula] vs. 12{+/-}2% of [Formula] for neuroticism, implicating the action of negative selection on low-frequency variants affecting gene regulation in the brain. Forward simulations confirmed that the ratio of low-frequency variant enrichment vs. common variant enrichment primarily depends on the mean selection coefficient of causal variants in the annotation, and can be used to predict the effect size variance of causal rare variants (MAF < 0.5%) in the annotation, informing their prioritization in whole-genome sequencing studies. Our results provide a deeper understanding of low-frequency variant functional architectures and guidelines for the design of association studies targeting functional classes of low-frequency and rare variants.

genetics

Unbiased construction of a temporally consistent morphological atlas of neonatal brain development

Premature birth increases the risk of developing neurocognitive and neurobe-havioural disorders. The mechanisms of altered brain development causing these disorders are yet unknown. Studying the morphology and function of the brain during maturation provides us not only with a better understanding of normal development, but may help us to identify causes of abnormal development and their consequences. A particular difficulty is to distinguish abnormal patterns of neurodevelopment from normal variation. The Developing Human Connectome Project (dHCP) seeks to create a detailed four-dimensional (4D) connectome of early life. This connectome may provide insights into normal as well as abnormal patterns of brain development. As part of this project, more than a thousand healthy fetal and neonatal brains will be scanned in vivo. This requires computational methods which scale well to larger data sets. We propose a novel groupwise method for the construction of a spatio-temporal model of mean morphology from cross-sectional brain scans at different gestational ages. This model scales linearly with the number of images and thus improves upon methods used to build existing public neonatal atlases, which derive correspondence between all pairs of images. By jointly estimating mean shape and longitudinal change, the atlas created with our method overcomes temporal inconsistencies, which are encountered when mean shape and intensity images are constructed separately for each time point. Using this approach, we have constructed a spatio-temporal atlas from 275 healthy neonates between 35 and 44 weeks post-menstrual age (PMA). The resulting atlas qualitatively preserves cortical details significantly better than publicly available atlases. This is moreover confirmed by a number of quantitative measures of the quality of the spatial normalisation and sharpness of the resulting template brain images.

neuroscience

Resolving the Full Spectrum of Human Genome Variation using Linked-Reads

Large-scale population based analyses coupled with advances in technology have demonstrated that the human genome is more diverse than originally thought. To date, this diversity has largely been uncovered using short read whole genome sequencing. However, standard short-read approaches, used primarily due to accuracy, throughput and costs, fail to give a complete picture of a genome. They struggle to identify large, balanced structural events, cannot access repetitive regions of the genome and fail to resolve the human genome into its two haplotypes. Here we describe an approach that retains long range information while harnessing the advantages of short reads. Starting from only [~]1ng of DNA, we produce barcoded short read libraries. The use of novel informatic approaches allows for the barcoded short reads to be associated with the long molecules of origin producing a novel datatype known as Linked-Reads. This approach allows for simultaneous detection of small and large variants from a single Linked-Read library. We have previously demonstrated the utility of whole genome Linked-Reads (lrWGS) for performing diploid, de novo assembly of individual genomes (Weisenfeld et al. 2017). In this manuscript, we show the advantages of Linked-Reads over standard short read approaches for reference based analysis. We demonstrate the ability of Linked-Reads to reconstruct megabase scale haplotypes and to recover parts of the genome that are typically inaccessible to short reads, including phenotypically important genes such as STRC, SMN1 and SMN2. We demonstrate the ability of both lrWGS and Linked-Read Whole Exome Sequencing (lrWES) to identify complex structural variations, including balanced events, single exon deletions, and single exon duplications. The data presented here show that Linked-Reads provide a scalable approach for comprehensive genome analysis that is not possible using short reads alone.

genomics

Leveraging polygenic functional enrichment to improve GWAS power

Functional genomics data has the potential to increase GWAS power by identifying SNPs that have a higher prior probability of association. Here, we introduce a method that leverages polygenic functional enrichment to incorporate coding, conserved, regulatory and LD-related genomic annotations into association analyses. We show via simulations with real genotypes that the method, Functionally Informed Novel Discovery Of Risk loci (FINDOR), correctly controls the false-positive rate at null loci and attains a 9-38% increase in the number of independent associations detected at causal loci, depending on trait polygenicity and sample size. We applied FINDOR to 27 independent complex traits and diseases from the interim UK Biobank release (average N=130K). Averaged across traits, we attained a 13% increase in genome-wide significant loci detected (including a 20% increase for disease traits) compared to un-weighted raw p-values that do not use functional data. We replicated the novel loci in independent UK Biobank and non-UK Biobank data, yielding a highly statistically significant replication slope (0.66-0.69) in each case. Finally, we applied FINDOR to the full UK Biobank release (average N=416K), attaining smaller relative improvements (consistent with simulations) but larger absolute improvements, detecting an additional 583 GWAS loci. In conclusion, leveraging functional enrichment using our method robustly increases GWAS power.

genetics

Detecting genome-wide directional effects of transcription factor binding on polygenic disease risk

Biological interpretation of GWAS data frequently involves analyzing unsigned genomic annotations comprising SNPs involved in a biological process and assessing enrichment for disease signal. However, it is often possible to generate signed annotations quantifying whether each SNP allele promotes or hinders a biological process, e.g., binding of a transcription factor (TF). Directional effects of such annotations on disease risk enable stronger statements about causal mechanisms of disease than enrichments of corresponding unsigned annotations. Here we introduce a new method, signed LD profile regression, for detecting such directional effects using GWAS summary statistics, and we apply the method using 382 signed annotations reflecting predicted TF binding. We show via theory and simulations that our method is well-powered and is well-calibrated even when TF binding sites co-localize with other enriched regulatory elements, which can confound unsigned enrichment methods. We further validate our method by showing that it recovers known transcriptional regulators when applied to molecular QTL in blood. We then apply our method to eQTL in 48 GTEx tissues, identifying 651 distinct TF-tissue expression associations at per-tissue FDR < 5%, including 30 associations with robust evidence of tissue specificity. Finally, we apply our method to 46 diseases and complex traits (average N = 289,617) and identify 77 annotation-trait associations at per-trait FDR < 5% representing 12 independent TF-trait associations, and we conduct gene-set enrichment analyses to characterize the underlying transcriptional programs. Our results implicate new causal disease genes (including causal genes at known GWAS loci), and in some cases suggest a detailed mechanism for a causal genes effect on disease. Our method provides a new way to leverage functional data to draw inferences about disease etiology.

genetics

Leveraging molecular QTL to understand the genetic architecture of diseases and complex traits

There is increasing evidence that many GWAS risk loci are molecular QTL for gene ex-pression (eQTL), histone modification (hQTL), splicing (sQTL), and/or DNA methylation (meQTL). Here, we introduce a new set of functional annotations based on causal posterior prob-abilities (CPP) of fine-mapped molecular cis-QTL, using data from the GTEx and BLUEPRINT consortia. We show that these annotations are very strongly enriched for disease heritability across 41 independent diseases and complex traits (average N = 320K): 5.84x for GTEx eQTL, and 5.44x for eQTL, 4.27-4.28x for hQTL (H3K27ac and H3K4me1), 3.61x for sQTL and 2.81x for meQTL in BLUEPRINT (all P [&le;] 1.39e-10), far higher than enrichments obtained using stan-dard functional annotations that include all significant molecular cis-QTL (1.17-1.80x). eQTL annotations that were obtained by meta-analyzing all 44 GTEx tissues generally performed best, but tissue-specific blood eQTL annotations produced stronger enrichments for autoimmune dis-eases and blood cell traits and tissue-specific brain eQTL annotations produced stronger enrich-ments for brain-related diseases and traits, despite high cis-genetic correlations of eQTL effect sizes across tissues. Notably, eQTL annotations restricted to loss-of-function intolerant genes from ExAC were even more strongly enriched for disease heritability (17.09x; vs. 5.84x for all genes; P = 4.90e-17 for difference). All molecular QTL except sQTL remained significantly enriched for disease heritability in a joint analysis conditioned on each other and on a broad set of functional annotations from previous studies, implying that each of these annotations is uniquely informative for disease and complex trait architectures.

genetics

Multimodal Surface Matching with Higher-Order Smoothness Constraints

The accurate alignment of brains is fundamental to the statistical sensitivity and spatial localisation of group studies in brain imaging, and cortical surface-based alignment is generally accepted to be superior to volume-based approaches at aligning cortical areas. However, human subjects have considerable variation in cortical folding, and in the location of cortical areas relative to these folds, which makes aligning cortical areas based on folding alone a challenging problem. The Multimodal Surface Matching (MSM) tool is a flexible spherical registration approach that enables accurate registration of surfaces based on a variety of different features. Using MSM, we have previously shown that using areal features such as resting state-networks and myelin maps to drive cross-subject surface alignment improves group task fMRI statistics and map sharpness. However, the initial implementation of MSMs regularisation function did not penalize all forms of surface distortion evenly. In some cases, this allowed peak distortions to exceed neurobiologically plausible limits unless the regularisation strength was increased, in which case this prevented the algorithm from fully maximizing surface alignment. Here, we propose a new regularisation penalty, derived from physically relevant equations of strain (deformation) energy, and demonstrate that its use leads to improved and more robust alignment of multi-modal imaging data. In addition, since spherical warps incorporate projection distortions that are unavoidable when mapping from a convoluted cortical surface to the sphere, we also propose constraints to enforce smooth deformation of cortical anatomies. We test the impact of this approach for longitudinal modeling of cortical development for neonates (born between 32 and 45 weeks) and demonstrate that the proposed method increases the biological interpretability of the distortion fields and improves the statistical significance of population-based analysis relative to other spherical methods.

neuroscience

Quantitative analysis of population-scale family trees using millions of relatives

Family trees have vast applications in multiple fields from genetics to anthropology and economics. However, the collection of extended family trees is tedious and usually relies on resources with limited geographical scope and complex data usage restrictions. Here, we collected 86 million profiles from publicly-available online data from genealogy enthusiasts. After extensive cleaning and validation, we obtained population-scale family trees, including a single pedigree of 13 million individuals. We leveraged the data to partition the genetic architecture of longevity by inspecting millions of relative pairs and to provide insights to population genetics theories on the dispersion of families. We also report a simple digital procedure to overlay other datasets with our resource in order to empower studies with population-scale genealogical data.\n\nOne Sentence SummaryUsing massive crowd-sourced genealogy data, we created a population-scale family tree resource for scientific studies.

genomics

Heritability enrichment of specifically expressed genes identifies disease-relevant tissues and cell types

Genetics can provide a systematic approach to discovering the tissues and cell types relevant for a complex disease or trait. Identifying these tissues and cell types is critical for following up on non-coding allelic function, developing ex-vivo models, and identifying therapeutic targets. Here, we analyze gene expression data from several sources, including the GTEx and PsychENCODE consortia, together with genome-wide association study (GWAS) summary statistics for 48 diseases and traits with an average sample size of 169,331, to identify disease-relevant tissues and cell types. We develop and apply an approach that uses stratified LD score regression to test whether disease heritability is enriched in regions surrounding genes with the highest specific expression in a given tissue. We detect tissue-specific enrichments at FDR < 5% for 34 diseases and traits across a broad range of tissues that recapitulate known biology. In our analysis of traits with observed central nervous system enrichment, we detect an enrichment of neurons over other brain cell types for several brain-related traits, enrichment of inhibitory over excitatory neurons for bipolar disorder but excitatory over inhibitory neurons for schizophrenia and body mass index, and enrichments in the cortex for schizophrenia and in the striatum for migraine. In our analysis of traits with observed immunological enrichment, we identify enrichments of T cells for asthma and eczema, B cells for primary biliary cirrhosis, and myeloid cells for Alzheimer's disease, which we validated with independent chromatin data. Our results demonstrate that our polygenic approach is a powerful way to leverage gene expression data for interpreting GWAS signal.

genetics