Search bioRxivSearch

Biology subjects

Xiao Wang

Publications and source records attributed to Xiao Wang.

5 recordsLinked to original sources

High-throughput Screening and CRISPR-Cas9 Modeling of Causal Lipid-associated Expression Quantitative Trait Locus Variants

Genome-wide association studies have identified a number of novel genetic loci linked to serum cholesterol and triglyceride levels. The causal DNA variants at these loci and the mechanisms by which they influence phenotype and disease risk remain largely unexplored. Expression quantitative trait locus analyses of patient liver and fat biopsies indicate that many lipid-associated variants influence gene expression in a cis-regulatory manner. However, linkage disequilibrium among neighboring SNPs at a genome-wide association study-implicated locus makes it challenging to pinpoint the actual variant underlying an association signal. We used a methodological framework for causal variant discovery that involves high-throughput identification of putative disease-causal loci through a functional reporter-based screen, the massively parallel reporter assay, followed by validation of prioritized variants in genome-edited human pluripotent stem cell models generated with CRISPR-Cas9. We complemented the stem cell models with CRISPR interference experiments in vitro and in knock-in mice in vivo. We provide validation for two high-priority SNPs, rs2277862 and rs10889356, being causal for lipid-associated expression quantitative trait loci. We also highlight the challenges inherent in modeling common genetic variation with these experimental approaches.\n\nAuthor SummaryGenome-wide association studies have identified numerous loci linked to a variety of clinical phenotypes. It remains a challenge to identify and validate the causal DNA variants in these loci. We describe the use of a high-throughput technique called the massively parallel reporter assay to analyze thousands of candidate causal DNA variants for their potential effects on gene expression. We use a combination of genome editing in human pluripotent stem cells, \"CRISPR interference\" experiments in other cultured human cell lines, and genetically modified mice to analyze the two highest-priority candidate DNA variants to emerge from the massively parallel reporter assay, and we confirm the relevance of the variants to nearby gene expression. These findings highlight a methodological framework with which to identify and functionally validate causal DNA variants.

Genomics

Simple multi-trait analysis identifies novel loci associated with growth and obesity measures

The ever-growing genome-wide association studies (GWAS) have revealed widespread pleiotropy. To exploit this, various methods which consider variant association with multiple traits jointly have been developed. However, most effort has been put on improving discovery power: how to replicate and interpret these discovered pleiotropic loci using multivariate methods has yet to be discussed fully. Using only multiple publicly available single-trait GWAS summary statistics, we develop a fast and flexible multi-trait framework that contains modules for (i) multi-trait genetic discovery, (ii) replication of locus pleiotropic profile, and (iii) multi-trait conditional analysis. The procedure is able to handle any level of sample overlap. As an empirical example, we discovered and replicated 23 novel pleiotropic loci for human anthropometry and evaluated their pleiotropic effects on other traits. By applying conditional multivariate analysis on the 23 loci, we discovered and replicated two additional multi-trait associated SNPs. Our results provide empirical evidence that multi-trait analysis allows detection of additional, replicable, highly pleiotropic genetic associations without genotyping additional individuals. The methods are implemented in a free and open source R package MultiABEL.\n\nAuthor summaryBy analyzing large-scale genomic data, geneticists have revealed widespread pleiotropy, i.e. single genetic variation can affect a wide range of complex traits. Methods have been developed to discover such genetic variants. However, we still lack insights into the relevant genetic architecture - What more can we learn from knowing the effects of these genetic variants?\n\nHere, we develop a fast and flexible statistical analysis procedure that includes discovery, replication, and interpretation of pleiotropic effects. The whole analysis pipeline only requires established genetic association study results. We also provide the mathematical theory behind the pleiotropic genetic effects testing.\n\nMost importantly, we show how a replication study can be essential to reveal new biology rather than solely increasing sample size in current genomic studies. For instance, we show that, using our proposed replication strategy, we can detect the difference in genetic effects between studies of different geographical origins.\n\nWe applied the method to the GIANT consortium anthropometric traits to discover new genetic associations, replicated in the UK Biobank, and provided important new insights into growth and obesity.\n\nOur pipeline is implemented in an open-source R package MultiABEL, sufficiently efficient that allows researchers to immediately apply on personal computers in minutes.

Genetics

Genome of octoploid plant maca (Lepidium meyenii) illuminates genomic basis for high altitude adaptation in the central Andes

Maca (Lepidium meyenii Walp, 2n = 8x = 64) of Brassicaceae family is an Andean economic plant cultivated on the 4000-4500 meters central sierra in Peru. Considering the rapid uplift of central Andes occurred 5 to 10 million years ago (Mya), an evolutionary question arises on how plants like maca acquire high altitude adaptation within short geological period. Here, we report the high-quality genome assembly of maca, in which two close-spaced maca-specific whole genome duplications (WGDs, [~] 6.7 Mya) were identified. Comparative genomics between maca and close-related Brassicaceae species revealed expansions of maca genes and gene families involved in abiotic stress response, hormone signaling pathway and secondary metabolite biosynthesis via WGDs. Retention and subsequent evolution of many duplicated genes may account for the morphological and physiological changes (i.e. small leaf shape and loss of vernalization) in maca for high altitude environment. Additionally, some duplicated maca genes under positive selection were identified with functions in morphological adaptation (i.e. MYB59) and development (i.e. GDPD5 and HDA9). Collectively, the octoploid maca genome sheds light on the important roles of WGDs in plant high altitude adaptation in the Andes.

Genomics

Ca-activation kinetics modulate successive puff/spark amplitude, duration and inter-event-interval correlations in a Langevin model of stochastic Ca release

Through theoretical analysis of the statistics of stochastic calcium (Ca2+) release (i.e., the amplitude, duration and inter-event interval of simulated Ca2+ puffs and sparks), we show that a Langevin description of the collective gating of Ca2+ channels may be a good approximation to the corresponding Markov chain model when the number of Ca2+ channels per Ca2+ release unit (CaRU) is in the physiological range. The Langevin description of stochastic Ca2+ release facilitates our investigation of correlations between successive puff/spark amplitudes, durations and inter-spark intervals, and how such puff/spark statistics depend on the number of channels per release site and the kinetics of Ca2+-mediated inactivation of open channels. When Ca2+ inactivation/de-inactivation rates are intermediate--i.e., the termination of Ca2+ puff/sparks is caused by the recruitment of inactivated channels--the correlation between successive puff/spark amplitudes is negative, while the correlations between puff/spark amplitudes and the duration of the preceding or subsequent inter-spark interval are positive. These correlations are significantly reduced when inactivation/deinactivation rates are extreme (slow or fast) and puff/sparks terminate via stochastic attrition.

Cell Biology

A population density and moment-based approach to modeling domain Ca-mediated inactivation of L-type Ca channels

We present a population density and moment-based description of the stochastic dynamics of domain Ca2+-mediated inactivation of L-type Ca2+ channels. Our approach accounts for the effect of heterogeneity of local Ca2+ signals on whole cell Ca2+ currents; however, in contrast with prior work, e.g., Sherman et al. (1990), we do not assume that Ca2+ domain formation and collapse are fast compared to channel gating. We demonstrate the population density and moment-based modeling approaches using a 12-state Markov chain model of an L-type Ca2+ channel introduced by Greenstein and Winslow (2002). Simulated whole cell voltage clamp responses yield an inactivation function for the whole cell Ca2+ current that agrees with the traditional approach when domain dynamics are fast. We analyze the voltage-dependence of Ca2+ inactivation that may occur via slow heterogeneous domains. Next, we find that when channel permeability is held constant, Ca2+-mediated inactivation of L-type channel increases as the domain time constant increases, because a slow domain collapse rate leads to increased mean domain [Ca2+] near open channels; conversely, when the maximum domain [Ca2+] is held constant, inactivation decreases as the domain time constant increases. Comparison of simulation results using population densities and moment equations confirms the computational efficiency of the moment-based approach, and enables the validation of two distinct methods of truncating and closing the open system of moment equations. In general, a slow domain time constant requires higher order moment truncation for agreement between moment-based and population density simulations.

Biophysics