Search bioRxiv⌕ Search

Biology subjects

Kelly, T. N.

Publications and source records attributed to Kelly, T. N..

2 recordsLinked to original sources

Whole genome sequence analysis of blood lipid levels in >66,000 individuals

Plasma lipids are heritable modifiable causal factors for coronary artery disease, the leading cause of death globally. Despite the well-described monogenic and polygenic bases of dyslipidemia, limitations remain in discovery of lipid-associated alleles using whole genome sequencing, partly due to limited sample sizes, ancestral diversity, and interpretation of potential clinical significance. Increasingly larger whole genome sequence datasets with plasma lipids coupled with methodologic advances enable us to more fully catalog the allelic spectrum for lipids. Here, among 66,329 ancestrally diverse (56% non-European ancestry) participants, we associate 428M variants from deep-coverage whole genome sequences with plasma lipids. Approximately 400M of these variants were not studied in prior lipids genetic analyses. We find multiple lipid-related genes strongly associated with plasma lipids through analysis of common and rare coding variants. We additionally discover several significantly associated rare non-coding variants largely at Mendelian lipid genes. Notably, we detect rare LDLR intronic variants associated with markedly increased LDL-C, similar to rare LDLR exonic variants. In conclusion, we conducted a systematic whole genome scan for plasma lipids expanding the alleles linked to lipids for multiple ancestries and characterize a clinically-relevant rare non-coding variant model for lipids.

genetics↗

A system for phenotype harmonization in the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program

Genotype-phenotype association studies often combine phenotype data from multiple studies to increase power. Harmonization of the data usually requires substantial effort due to heterogeneity in phenotype definitions, study design, data collection procedures, and data set organization. Here we describe a centralized system for phenotype harmonization that includes input from phenotype domain and study experts, quality control, documentation, reproducible results, and data sharing mechanisms. This system was developed for the National Heart, Lung and Blood Institutes Trans-Omics for Precision Medicine (TOPMed) program, which is generating genomic and other omics data for >80 studies with extensive phenotype data. To date, 63 phenotypes have been harmonized across thousands of participants from up to 17 TOPMed studies per phenotype. We discuss the challenges faced in this undertaking and how they were addressed. The harmonized phenotype data and associated documentation have been submitted to National Institutes of Health data repositories for controlled-access by the scientific community. We also provide materials to facilitate future harmonization efforts by the community, which include (1) the code used to generate the 63 harmonized phenotypes, enabling others to reproduce, modify or extend these harmonizations to additional studies; and (2) results of labeling thousands of phenotype variables with controlled vocabulary terms.

genetics↗