bioRxiv · 10.1101/2020.07.09.196535
Improved analyses of GWAS summary statistics by reducing data heterogeneity and errors
Abstract
Summary statistics from genome-wide association studies (GWAS) have facilitated the development of various summary data-based methods, which typically require a reference sample for linkage disequilibrium (LD) estimation. Analyses using these methods may be biased by errors in GWAS summary data and heterogeneity between GWAS and LD reference. Here we propose a quality control method, DENTIST, that leverages LD among genetic variants to detect and eliminate errors in GWAS or LD reference and heterogeneity between the two. Through simulations, we demonstrate that DENTIST substantially reduces false-positive rate (FPR) in detecting secondary signals in the summary-data-based conditional and joint (COJO) association analysis, especially for imputed rare variants (FPR reduced from >28% to <2% in the presence of heterogeneity between GWAS and LD reference). We further show that DENTIST can improve other summary-data-based analyses such as fine-mapping analysis, and integrative analysis of GWAS and expression quantitative trait locus data.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenhan Chen, Yang Wu, Zhili Zheng, Ting Qi, Peter M Visscher, Zhihong Zhu, Jian Yang. 2020-07-12. Improved analyses of GWAS summary statistics by reducing data heterogeneity and errors. https://doi.org/10.1101/2020.07.09.196535
Cite the original work for its findings. Save a collection to share your selection of sources.