Search bioRxivSearch

Biology subjects

Cho, M. H.

Publications and source records attributed to Cho, M. H..

4 recordsLinked to original sources

Efficient variant set mixed model association tests for continuous and binary traits in large-scale whole genome sequencing studies

With advances in Whole Genome Sequencing (WGS) technology, more advanced statistical methods for testing genetic association with rare variants are being developed. Methods in which variants are grouped for analysis are also known as variant-set, gene-based, and aggregate unit tests. The burden test and Sequence Kernel Association Test (SKAT) are two widely used variant-set tests, which were originally developed for samples of unrelated individuals and later have been extended to family data with known pedigree structures. However, computationally-efficient and powerful variant-set tests are needed to make analyses tractable in large-scale WGS studies with complex study samples. In this paper, we propose the variant-Set Mixed Model Association Tests (SMMAT) for continuous and binary traits using the generalized linear mixed model framework. These tests can be applied to large-scale WGS studies involving samples with population structure and relatedness, such as in the National Heart, Lung, and Blood Institutes Trans-Omics for Precision Medicine (TOPMed) program. SMMAT tests share the same null model for different variant sets, and a virtue of this null model, which includes covariates only, is that it needs to be only fit once for all tests in each genome-wide analysis. Simulation studies show that all the proposed SMMAT tests correctly control type I error rates for both continuous and binary traits in the presence of population structure and relatedness. We also illustrate our tests in a real data example of analysis of plasma fibrinogen levels in the TOPMed program (n = 23,763), using the Analysis Commons, a cloud-based computing platform.

genetics

Genotype Imputation Performance of Three Reference Panels Using African Ancestry Individuals

Genotype imputation is used to estimate unobserved genotypes from genome-wide maker data, to increase genome coverage and power for genome-wide association studies. Imputation has been most successful for European ancestry populations in which very large reference panels are available. Smaller subsets of African descent populations are available in 1000 Genomes (1000G), the Consortium on Asthma among African-Ancestry Populations in the Americas (CAAPA) and the Haplotype Reference Consortium (HRC). We aimed to compare the performance of these reference panels when imputing variation in 3,747 African Americans (AA) from 2 cohorts (HCV and COPDGene) genotyped using the Illumina Omni family of microarrays. The haplotypes of 2,504 individuals (from 1000G), 883 (from CAAPA) and 32,611 (from HRC) were used as reference. We compared the performance of these panels based on number of variants, imputation quality, imputation accuracy and coverage. In both cohorts, 1000G imputed 1.5-1.6x more variants compared to CAAPA and 1.2x more variants than HRC. Similar findings were observed for variants with higher imputation quality (R2>0.5) and for rare, low frequency, and common variants. When merging the results of the three panels the total number of imputed variants was 62M-63M with 20M overlapping variants imputed by all three panels, and a range of 5 to 15M unique variants imputed exclusively with one of the three panels. For overlapping variants, imputation quality was highest for HRC, followed by 1000G, then CAAPA, and improved as the minor allele frequency increased. The 1000G, HRC and CAAPA participants of African ancestry provided high performance and accuracy for imputation of African American admixed individuals, increasing the total number of variants with high quality available for subsequent analyses. These three panels are complementary and would benefit from the development of an integrated African reference panel, including data from multiple sources and populations.

genetics

Leveraging lung tissue transcriptome to uncover candidate causal genes in COPD genetic associations

We collated 129 non-overlapping risk loci for chronic obstructive pulmonary disease (COPD) from the GWAS literature. Using recent and complementary integrative genomics approaches, combining GWAS and lung eQTL results, we identified 12 novel COPD loci and corresponding causal genes. In addition, we mapped candidate causal genes for 60 out of the 129 GWAS-nominated loci as well as for four sub-genome-wide significant COPD risk loci derived from the largest GWAS on COPD. Mapping causal genes in lung tissue represents an important contribution on the genetics of COPD, enriches our biological interpretation of GWAS findings, and brings us closer to clinical translation of genetic associations.

genomics

Integrative genomics analysis identifies ACVR1B as a candidate causal gene of emphysema distribution in non-alpha 1-antitrypsin deficient smokers

BackgroundSeveral genetic risk loci associated with emphysema apico-basal distribution (EABD) have been identified through genome-wide association studies (GWAS), but the biological functions of these variants are unknown. To characterize gene regulatory functions of EABD-associated variants, we integrated EABD GWAS results with 1) a multi-tissue panel of expression quantitative trait loci (eQTL) from subjects with COPD and the GTEx project and 2) epigenomic marks from 127 cell types in the Roadmap Epigenomics project. Functional validation was performed for a variant near ACVR1B.\n\nResultsSNPs from 168 loci with P-values<5x10-5 in the largest GWAS meta-analysis of EABD (Boueiz A. et al, AJRCCM 2017) were analyzed. 54 loci overlapped eQTL regions from our multi-tissue panel, and 7 of these loci showed a high probability of harboring a single, shared GWAS and eQTL causal variant (colocalization posterior probability[&ge;]0.9). 17 cell types exhibited greater than expected overlap between EABD loci and DNase-I hypersensitive peaks, DNaseI hotspots, enhancer marks, or digital DNaseI footprints (permutation P-value < 0.05), with the strongest enrichment observed in CD4+, CD8+, and regulatory T cells. A region near ACVR1B demonstrated significant colocalization with a lung eQTL and overlapped DNase-I hypersensitive regions in multiple cell types, and reporter assays in human bronchial epithelial cells confirmed allele-specific regulatory activity for the lead variant, rs7962469.\n\nConclusionsIntegrative analysis highlights candidate causal genes, regulatory variants, and cell types that may contribute to the pathogenesis of emphysema distribution. These findings will enable more accurate functional validation studies and better understanding of emphysema distribution biology.

genomics