Search bioRxivSearch

Biology subjects

Carey, D. J.

Publications and source records attributed to Carey, D. J..

4 recordsLinked to original sources

Rare Variant Pathogenicity Triage and Inclusion of Synonymous Variants Improves Analysis of Disease Associations

Many G protein-coupled receptors (GPCRs) lack common variants that lead to reproducible genome-wide disease associations. Here we used rare variant approaches to assess the disease associations of 85 orphan or understudied GPCRs in an unselected cohort of 51,289 individuals. Rare loss-of-function variants, missense variants predicted to be pathogenic or likely pathogenic, and a subset of rare synonymous variants were used as independent data sets for sequence kernel association testing (SKAT). Strong, phenome-wide disease associations shared by two or more variant categories were found for 39% of the GPCRs. Validating the bioinformatics and SKAT analyses, functional characterization of rare missense and synonymous variants of GPR39, a Family A GPCR, showed altered expression and/or Zn2+-mediated signaling for members of both variant classes. Results support the utility of rare variant analyses for identifying disease associations for genes that lack common variants, while also highlighting the functional importance of rare synonymous variants.\n\nAuthor summaryRare variant approaches have emerged as a viable way to identify disease associations for genes without clinically important common variants. Rare synonymous variants are generally considered benign. We demonstrate that rare synonymous variants represent a potentially important dataset for deriving disease associations, here applied to analysis of a set of orphan or understudied GPCRs. Synonymous variants yielded disease associations in common with loss-of-function or missense variants in the same gene. We rationalize their associations with disease by confirming their impact on expression and agonist activation of a representative example, GPR39. This study highlights the importance of rare synonymous variants in human physiology, and argues for their routine inclusion in any comprehensive analysis of genomic variants as potential causes of disease.

genetics

Profiling and leveraging relatedness in a precision medicine cohort of 92,455 exomes

Large-scale human genetics studies are ascertaining increasing proportions of populations as they continue growing in both number and scale. As a result, the amount of cryptic relatedness within these study cohorts is growing rapidly and has significant implications on downstream analyses. We demonstrate this growth empirically among the first 92,455 exomes from the DiscovEHR cohort and, via a custom simulation framework we developed called SimProgeny, show that these measures are in-line with expectations given the underlying population and ascertainment approach. For example, we identified [~]66,000 close (first- and second-degree) relationships within DiscovEHR involving 55.6% of study participants. Our simulation results project that >70% of the cohort will be involved in these close relationships as DiscovEHR scales to 250,000 recruited individuals. We reconstructed 12,574 pedigrees using these relationships (including 2,192 nuclear families) and leveraged them for multiple applications. The pedigrees substantially improved the phasing accuracy of 20,947 rare, deleterious compound heterozygous mutations. Reconstructed nuclear families were critical for identifying 3,415 de novo mutations in [~]1,783 genes. Finally, we demonstrate the segregation of known and suspected disease-causing mutations through reconstructed pedigrees, including a tandem duplication in LDLR causing familial hypercholesterolemia. In summary, this work highlights the prevalence of cryptic relatedness expected among large healthcare population genomic studies and demonstrates several analyses that are uniquely enabled by large amounts of cryptic relatedness.

genomics

A Model for Genome-First Care: Returning Secondary Genomic Findings to Participants and Their Healthcare Providers in a Large Research Cohort

BackgroundResearch cohorts with linked genomic data exist, or are being developed, at many research centers. Within any such \"sequenced cohort\" of more than 100 participants, it is likely that there are participants with previously undisclosed risk for life-threatening monogenic diseases that could be identified with targeted analysis of their existing data. Identification of such disease-associated findings are not usually primary to the enrollment research goals. At Geisinger Health System, MyCode(R) Community Health Initiative (MyCode) participants represent one such large sequenced cohort. Since 2013, MyCode participants in discovery research have been consented for secondary analysis of their existing research genomic sequences to allow delivery of medically actionable findings to them and their healthcare providers. This return of genomic results program was developed to manage an anticipated 3.5% of MyCode participants who will receive clinically confirmed genomic variants from an approved gene list out of more than 150,000 total participants. Risk-associated DNA sequences alone without any clinical parameter, prompt \"genome-first\" follow-up encounters.\n\nMethodsThis article describes our process for generating clinical grade results from research-based genomic sequencing data, delivering results to patients and their providers, facilitating targeted clinical evaluations of patients and promoting cascade testing of at-risk relatives. We also summarize our early data about the results generated during this process and our ability to contact patients and their providers to disclose the information.\n\nResultsThis process has been used to generate 343 results on 339 patients. 93% of patients with a result have been successfully contacted about their results as evidenced by direct interaction about their result with the research team or a healthcare provider. 222 healthcare providers have been notified of a result on one or more patient through this result delivery process.\n\nConclusionsHere we describe the existing GHS model to deliver genomic data into the electronic medical record and the clinical interactions that are prompted and supported. Elements of this genome-first care model can be applied in other healthcare settings and in national efforts, such as \"All of Us\", that wish to establish programs for returning genomic results to research participants.

genomics

Profiling copy number variation and disease associations from 50,726 DiscovEHR Study exomes

Copy number variants (CNVs) are a substantial source of genomic variation and contribute to a wide range of human disorders. Gene-disrupting exonic CNVs have important clinical implications as they can underlie variability in disease presentation and susceptibility. The relationship between exonic CNVs and clinical traits has not been broadly explored at the population level, primarily due to technical challenges. We surveyed common and rare CNVs in the exome sequences of 50,726 adult DiscovEHR study participants with linked electronic health records (EHRs). We evaluated the diagnostic yield and clinical expressivity of known pathogenic CNVs, and performed tests of association with EHR-derived serum lipids, thereby evaluating the relationship between CNVs and complex traits and phenotypes in an unbiased, real-world clinical context. We identified CNVs from megabase to exon-level resolution, demonstrating reliable, high-throughput detection of clinically relevant exonic CNVs. In doing so, we created a catalog of high-confidence common and rare CNVs and refined population frequency estimates of known and novel gene-disrupting CNVs. Our survey among an unselected clinical population provides further evidence that neuropathy-associated duplications and deletions in 17p12 have similar population prevalence but are clinically under-diagnosed. Similarly, adults who harbor 22q11.2 deletions frequently had EHR documentation of neurodevelopmental/neuropsychiatric disorders and congenital anomalies, but not a formal genetic diagnosis (i.e., deletion). In an exome-wide association study of lipid levels, we identified a novel five-exon duplication within LDLR segregating in a large kindred with features of familial hypercholesterolemia. Exonic CNVs provide new opportunities to understand and diagnose human disease.

genomics