Search bioRxivSearch

Biology subjects

Qi, G.

Publications and source records attributed to Qi, G..

5 recordsLinked to original sources

Mendelian Randomization Analysis Using Mixture Models (MRMix) for Genetic Effect-Size-Distribution Leads to Robust Estimation of Causal Effects

We propose a novel method for robust estimation of causal effects in two-sample Mendelian randomization analysis using potentially large number of genetic instruments. We consider a \"working model\" for bi-variate effect-size distribution across pairs of traits in the form of normal-mixtures which assumes existence of a fraction of the genetic markers that are valid instruments, i.e. they have only direct effect on one trait, while other markers can have potentially correlated, direct and indirect effects, or have no effects at all. We show that model motivates a simple method for estimating causal effect ({theta}) through a procedure for maximizing the probability concentration of the residuals, [Formula], at the \"null\" component of a two-component normal-mixture model. Simulation studies showed that MRMix provides nearly unbiased or/and substantially more robust estimates of causal effects compared to alternative methods under various scenarios. Further, the studies showed that MRMix is sensitive to direction and can achieve much higher efficiency (up to 3-4 fold) relative to other comparably robust estimators. We applied the proposed methods for conducting MR analysis using largest publicly available datasets across a number of risk-factors and health outcomes. Notable findings included identification of causal effects of genetically determined BMI and ageat-menarche, which have relationship among themselves, on the risk of breast cancer; detrimental effect of HDL on the risk of breast cancer; no causal effect of HDL and triglycerides on the risk of coronary artery disease; a strong detrimental effect of BMI, but no causal effect of years of education, on the risk of major depressive disorder.

genetics

Exercise-related genes analysis of Mongolian Horse - Abaga horse and Wushen horse

The Mongolian horses, as a neglected scientific resource, have excellent endurance and stress resistance to adapt to the cold and harsh plateau conditions. Intraspecific genetic diversity is mainly embodied in various genetic advantages of different branches of Mongolian horse. Abaga horse is better than Wushen horse in running speed, for example. Because people pay progressively attention to the athletic performance of horse, such as horse racing in Mongolias Naadam festival, we expect to guide the exercise-oriented breeding of horses through genomics research. We obtained the clean data of 630,535,376,400 bp through the entire genome second-generation sequencing for the whole blood of 4 Abaga horses and 10 Wushen horses. Based on the data analysis of single nucleotide polymorphism (SNP), we severally detected that 479 and 943 positively selected genes, particularly exercise-related, were mainly enriched on equine chromosome 4 in Abaga horses and Wushen horses, which implied that the chromosome 4 may be associated with the evolution of the Mongolian horse and athletic performance. Four hundred and forty genes of positive selection were enriched in 12 exercise-related pathways and narrowed in 21 exercise-related genes in Abaga horse, which were distinguished from Wushen horse. So, we speculated that the Abaga horse may have oriented genes for the motorial mechanism and 21 exercise-related genes also provided molecular genetic basis for exercise-directed breeding of Mongolian horse.

genomics

Effect Size and Power in fMRI Group Analysis

Multi-subject functional magnetic resonance imaging (fMRI) analysis is often concerned with determining whether there exists a significant population-wide activation in a comparison between two or more conditions. Typically this is assessed by testing the average value of a contrast of parameter estimates (COPE) against zero in a general linear model (GLM) analysis. In this work we investigate several aspects of this type of analysis. First, we study the effects of sample size on the sensitivity and reliability of the group analysis, allowing us to evaluate the ability of small sampled studies to effectively capture population-level effects of interest. Second, we assess the difference in sensitivity and reliability when using volumetric or surface based data. Third, we investigate potential biases in estimating effect sizes as a function of sample size. To perform this analysis we utilize the task-based fMRI data from the 500-subject release from the Human Connectome Project (HCP). We treat the complete collection of subjects (N = 491) as our population of interest, and perform a single-subject analysis on each subject in the population. We investigate the ability to recover population level effects using a subset of the population and standard analytical techniques. Our study shows that sample sizes of 40 are generally able to detect regions with high effect sizes (Cohens d > 0.8), while sample sizes closer to 80 are required to reliably recover regions with medium effect sizes (0.5 < d < 0.8). We find little difference in results when using volumetric or surface based data with respect to standard mass-univariate group analysis. Finally, we conclude that special care is needed when estimating effect sizes, particularly for small sample sizes.

neuroscience

Heritability Informed Power Optimization (HIPO) Leads to Enhanced Detection of Genetic Associations Across Multiple Traits

Genome-wide association studies have shown that pleiotropy is a common phenomenon that can potentially be exploited for enhanced detection of susceptibility loci. We propose heritability informed power optimization (HIPO) for conducting powerful pleiotropic analysis using summary-level association statistics. We find optimal linear combinations of association coefficients across traits that are expected to maximize non-centrality parameter for the underlying test statistics, taking into account estimates of heritability, sample size variations and overlaps across the traits. Simulation studies show that the proposed method has correct type I error, robust to population stratification and leads to desired genome-wide enrichment of association signals. Application of the proposed method to publicly available data for three groups of genetically related traits, lipids (N=188,577), psychiatric diseases (Ncase=33,332, Ncontrol=27,888) and social science traits (N ranging between 161,460 to 298,420 across individual traits) increased the number of genome-wide significant loci by 12%, 200% and 50%, respectively, compared to those found by analysis of individual traits. Evidence of replication is present for many of these loci in subsequent larger studies for individual traits. HIPO can potentially be extended to high-dimensional phenotypes as a way of dimension reduction to maximize power for subsequent genetic association testing.

genetics

Estimation of complex effect-size distributions using summary-level statistics from genome-wide association studies across 32 complex traits and implications for the future

Summary-level statistics from genome-wide association studies are now widely used to estimate heritability and co-heritability of traits using the popular linkage-disequilibrium-score (LD-score) regression method. We develop a likelihood-based approach for analyzing summary-level statistics and external LD information to estimate common variants effect-size distributions, characterized by proportion of underlying susceptibility SNPs and a flexible normal-mixture model for their effects. Analysis of summary-level results across 32 GWAS reveals that while all traits are highly polygenic, there is wide diversity in the degrees of polygenicity. The effect-size distributions for susceptibility SNPs could be adequately modeled by a single normal distribution for traits related to mental health and ability and by a mixture of two normal distributions for all other traits. Among quantitative traits, we predict the sample sizes needed to identify SNPs which explain 80% of GWAS heritability to be between 300K-500K for some of the early growth traits, between 1-2 million for some anthropometric and cholesterol traits and multiple millions for body mass index and some others. The corresponding predictions for disease traits are between 200K-400K for inflammatory bowel diseases, close to one million for a variety of adult onset chronic diseases and between 1-2 million for psychiatric diseases.

genetics