Search bioRxivSearch

Biology subjects

Liu, D. J.

Publications and source records attributed to Liu, D. J..

3 recordsLinked to original sources

Proper Conditional Analysis in the Presence of Missing Data Identified Novel Independently Associated Low Frequency Variants in Nicotine Dependence Genes

Meta-analysis of genetic association studies increases sample size and the power for mapping complex traits. Existing methods are mostly developed for datasets without missing values. In practice, genotype imputation is not always effective, e.g. when targeted genotyping/sequencing assays are used or when the un-typed genetic variant is rare. Therefore, contributed summary statistics often contain missing values. Naive extensions of existing methods either replace missing summary statistics with 0 or discard studies with missing data. These approaches can bias genetic effect estimates and lead to seriously inflated type-I or II errors in conditional analysis, which is a critical tool for identifying independently associated variants.\n\nTo address this challenge and complement imputation methods, we developed a method to combine summary statistics across participating studies and consistently estimate joint effects, even when the contributed summary statistics contain large amount of missing values. Based on this estimator, we propose a score statistic we call PCBS (partial correlation based score statistic) for conditional analysis of single-variant and gene-level associations. Through extensive analysis of simulated and real data, we showed that the new method produces well-calibrated type-I errors and is substantially more powerful than existing approaches. We applied the proposed approach to analyze the CHRNA5-CHRNB4-CHRNA3 locus in a large-scale meta-analysis for cigarettes-per-day. Using the new method, we identified three novel variants, independent of known association signals, which were otherwise missed by alternative methods. Together, the phenotypic variance explained by these variants is .46%, improving that of previously reported associations by 17%. These findings illustrate the extent of locus allelic heterogeneity and can help pinpoint causal variants.\n\nAUTHOR SUMMARYIt is of great interest to estimate the joint and conditional effects of multiple correlated variants from large scale meta-analysis, in order to fine map causal variants and understand the genetic architecture for complex traits. The contributed summary statistics from participating studies in a meta-analysis often contain missing values, as the imputation methods are not often effective, especially when the underlying genetic variant is rare or the participating studies use targeted genotyping array that is not suitable for imputation. Existing meta-analysis methods do not properly handle missing data, and can incorrectly estimate correlations between score statistics. As a result, they can produce highly biased estimates of joint effects and highly inflated type-I errors for conditional analysis, which will in turn result in overestimated phenotypic variance explained and incorrect identification of causal variants. We systematically evaluated this bias and proposed a novel partial correlation based score statistic. The new statistic has valid type-I errors for conditional analysis and much higher power than the existing methods, even when the contributed summary statistics in the meta-analysis contain a large fraction of missing values. We expect this method to be highly useful in the sequencing age for complex trait genetics.

genetics

Association Analysis and Meta-Analysis of Multi-allelic Variants for Large Scale Sequence Data

MotivationThere is great interest to understand the impact of rare variants in human diseases using large sequence datasets. In deep sequences datasets of >10,000 samples, [~]10% of the variant sites are observed to be multi-allelic. Many of the multi-allelic variants have been shown to be functional and disease relevant. Proper analysis of multi-allelic variants is critical to the success of a sequencing study, but existing methods do not properly handle multi-allelic variants and can produce highly misleading association results.\n\nResultsWe propose novel methods to encode multi-allelic sites, conduct single variant and gene-level association analyses, and perform meta-analysis for multi-allelic variants. We evaluated these methods through extensive simulations and the study of a large meta-analysis of [~]18,000 samples on the cigarettes-per-day phenotype. We showed that our joint modeling approach provided an unbiased estimate of genetic effects, greatly improved the power of single variant association tests, and enhanced gene-level tests over existing approaches.\n\nAvailabilitySoftware packages implementing these methods are available at (https://github.com/zhanxw/rvtests http://genome.sph.umich.edu/wiki/RareMETAL).\n\nContactxiaowei.zhan@utsouthwestem.edu; dajiang.liu@psu.edu

bioinformatics

Exome chip meta-analysis elucidates the genetic architecture of rare coding variants in smoking and drinking behavior

BackgroundSmoking and alcohol use behaviors in humans have been associated with common genetic variants within multiple genomic loci. Investigation of rare variation within these loci holds promise for identifying causal variants impacting biological mechanisms in the etiology of disordered behavior. Microarrays have been designed to genotype rare nonsynonymous and putative loss of function variants. Such variants are expected to have greater deleterious consequences on gene function than other variants, and significantly contribute to disease risk.\n\nMethodsIn the present study, we analyzed [~]250,000 rare variants from 17 independent studies. Each variant was tested for association with five addiction-related phenotypes: cigarettes per day, pack years, smoking initiation, age of smoking initiation, and alcoholic drinks per week. We conducted single variant tests of all variants, and gene-based burden tests of nonsynonymous or putative loss of function variants with minor allele frequency less than 1%.\n\nResultsMeta-analytic sample sizes ranged from 70,847 to 164,142 individuals, depending on the phenotype. Known loci tagged by common variants replicated, but there was no robust evidence for individually associated rare variants, either in gene based or single variant tests. Using a modified method-of-moment approach, we found that all low frequency coding variants, in aggregate, contributed 1.7% to 3.6% of the phenotypic variation for the five traits (p<.05).\n\nConclusionsThe findings indicate that rare coding variants contribute to phenotypic variation, but that much larger samples and/or denser genotyping of rare variants will be required to successfully identify associations with these phenotypes, whether individual variants or gene- based associations.

genetics