Search bioRxivSearch

Biology subjects

Li Wang

Publications and source records attributed to Li Wang.

4 recordsLinked to original sources

High-throughput Screening and CRISPR-Cas9 Modeling of Causal Lipid-associated Expression Quantitative Trait Locus Variants

Genome-wide association studies have identified a number of novel genetic loci linked to serum cholesterol and triglyceride levels. The causal DNA variants at these loci and the mechanisms by which they influence phenotype and disease risk remain largely unexplored. Expression quantitative trait locus analyses of patient liver and fat biopsies indicate that many lipid-associated variants influence gene expression in a cis-regulatory manner. However, linkage disequilibrium among neighboring SNPs at a genome-wide association study-implicated locus makes it challenging to pinpoint the actual variant underlying an association signal. We used a methodological framework for causal variant discovery that involves high-throughput identification of putative disease-causal loci through a functional reporter-based screen, the massively parallel reporter assay, followed by validation of prioritized variants in genome-edited human pluripotent stem cell models generated with CRISPR-Cas9. We complemented the stem cell models with CRISPR interference experiments in vitro and in knock-in mice in vivo. We provide validation for two high-priority SNPs, rs2277862 and rs10889356, being causal for lipid-associated expression quantitative trait loci. We also highlight the challenges inherent in modeling common genetic variation with these experimental approaches.\n\nAuthor SummaryGenome-wide association studies have identified numerous loci linked to a variety of clinical phenotypes. It remains a challenge to identify and validate the causal DNA variants in these loci. We describe the use of a high-throughput technique called the massively parallel reporter assay to analyze thousands of candidate causal DNA variants for their potential effects on gene expression. We use a combination of genome editing in human pluripotent stem cells, \"CRISPR interference\" experiments in other cultured human cell lines, and genetically modified mice to analyze the two highest-priority candidate DNA variants to emerge from the massively parallel reporter assay, and we confirm the relevance of the variants to nearby gene expression. These findings highlight a methodological framework with which to identify and functionally validate causal DNA variants.

Genomics

On the Origin and Evolutionary Consequences of Gene Body DNA Methylation

In plants, CG DNA methylation is prevalent in the transcribed regions of many constitutively expressed genes (\"gene body methylation; gbM\"), but the origin and function of gbM remain unknown. Here we report the discovery that Eutrema salsugineum has lost gbM from its genome, the first known instance for an angiosperm. Of all known DNA methyltransferases, only CHROMOMETHYLASE 3 (CMT3) is missing from E. salsugineum. Identification of an additional angiosperm, Conringia planisiliqua, which independently lost CMT3 and gbM supports that CMT3 is required for the establishment of gbM. Detailed analyses of gene expression, the histone variant H2A.Z and various histone modifications in E. salsugineum and in Arabidopsis thaliana epiRILs found no evidence in support of any role for gbM in regulating transcription or affecting the composition and modifications of chromatin over evolutionary time scales.

Plant Biology

Recent demography drives changes in linked selection across the maize genome

Genetic diversity is shaped by the interaction of drift and selection, but the details of this interaction are not well understood. The impact of genetic drift in a population is largely determined by its demographic history, typically summarized by its long-term effective population size (Ne). Rapidly changing population demographics complicate this relationship, however. To better understand how changing demography impacts selection, we used whole-genome sequencing data to investigate patterns of linked selection in domesticated and wild maize (teosinte). We produce the first whole-genome estimate of the demography of maize domestication, showing that maize was reduced to approximately 5% the population size of teosinte before it experienced rapid expansion post-domestication to population sizes much larger than its ancestor. Evaluation of patterns of nucleotide diversity in and near genes shows little evidence of selection on beneficial amino acid substitutions, and that the domestication bottleneck led to a decline in the efficiency of purifying selection in maize. Young alleles, however, show evidence of much stronger purifying selection in maize, reflecting the much larger effective size of present day populations. Our results demonstrate that recent demographic change -- a hallmark of many species including both humans and crops -- can have immediate and wide-ranging impacts on diversity that conflict with would-be expectations based on Ne alone.

Evolutionary Biology

Comprehensive mutational scanning of a kinase in vivo reveals substrate-dependent fitness landscapes

Deep mutational scanning has emerged as a promising tool for mapping sequence-activity relationships in proteins1-4, RNA5 and DNA6-8. In this approach, diverse variants of a sequence of interest are first ranked according to their activities in a relevant pooled assay, and this ranking is then used to infer the shape of the fitness landscape around the wild-type sequence. Little is currently know, however, about the degree to which such fitness landscapes are dependent on the specific assay conditions from which they are inferred. To explore this issue, we performed deep mutational scanning of APH(3)II, a Tn5 transposon-derived kinase that confers resistance to aminoglycoside antibiotics9, in E. coli under selection with each of six structurally diverse antibiotics at a range of inhibitory concentrations. We found that the resulting fitness landscapes showed significant dependence on both antibiotic structure and concentration. This shows that the notion of essential amino acid residues is context-dependent, but also that this dependence can be exploited to guide protein engineering. Specifically, we found that differential analysis of fitness landscapes allowed us to generate synthetic APH(3)II variants with orthogonal substrate specificities.

Molecular Biology