Search bioRxivSearch

Biology subjects

Wei Wang

Publications and source records attributed to Wei Wang.

9 recordsLinked to original sources

Assessment of single cell RNA-seq normalization methods

We have assessed the performance of seven normalization methods for single cell RNA-seq using data generated from dilution of RNA samples. Our analyses showed that methods considering spike-in ERCC RNA molecules significantly outperformed those not considering ERCCs. This work provides a guidance of selecting normalization methods to remove technical noise in single cell RNA-seq data.

Bioinformatics

Optimizing multiplex CRISPR/Cas9-based genome editing for wheat

BackgroundCRISPR/Cas9-based genome editing holds great promise to accelerate the development of new crop varieties by providing a powerful tool to modify the genomic regions controlling major agronomic traits. To diversify the set of tools available for wheat genome engineering, we have established a tRNA-based multiplex gene editing strategy for hexaploid wheat.\n\nResultsThe functionality of the various CRISPR/Cas9 components was assessed using the transient expression in the wheat protoplasts followed by next-generation sequencing (NGS) of the targeted genomic regions. The efficiency of wheat codon-optimized Cas9 for targeted gene editing in wheat was validated. Multiple single guide RNAs (gRNAs) were evaluated for the ability to edit the homoeologous copies of four genes affecting some important agronomic traits in wheat. Low correspondence was found between the gRNA efficiency predicted bioinformatically and that assessed in the transient expression assay. A multiplex gene editing construct with several gRNA-tRNA units under the control of a single promoter for the RNA polymerase III generated indels at the targets sites with the efficiency comparable to that obtained for a single gRNA construct.\n\nConclusionsBy integrating the protoplast transformation assay with multiplexed NGS, it is possible to perform fast functional screens for a large number of gRNAs and to optimize constructs for effective editing of multiple independent targets in the wheat genome. The multiplexing capacity of the tandemly arrayed tRNA-gRNA construct is well suited for the simultaneous editing of the redundant gene copies in the allopolyploid genomes or genomic regions beneficially affecting multiple agronomic traits. A polycistronic gene construct that can be quickly assembled using the Golden Gate reaction along with the wheat codon optimized Cas9 will further expand the set of tools available for engineering the wheat genome.

Genomics

Assessing the measurement transfer function of single-cell RNA sequencing

Recently, measurement of RNA at single cell resolution has yielded surprising insights. Methods for single-cell RNA sequencing (scRNA-seq) have received considerable attention, but the broad reliability of single cell methods and the factors governing their performance are still poorly known. Here, we conducted a large-scale control experiment to assess the transfer function of three scRNA-seq methods and factors modulating the function. All three methods detected greater than 70% of the expected number of genes and had a 50% probability of detecting genes with abundance greater than 2 to 4 molecules. Despite the small number of molecules, sequencing depth significantly affected gene detection. While biases in detection and quantification were qualitatively similar across methods, the degree of bias differed, consistent with differences in molecular protocol. Measurement reliability increased with expression level for all methods and we conservatively estimate the measurement transfer functions to be linear above ~5-10 molecules. Based on these extensive control studies, we propose that RNA-seq of single cells has come of age, yielding quantitative biological information.

Genomics

Finding De novo methylated DNA motifs

Increasing evidence has shown that posttranslational modifications (PTMs) such as methylation and hydroxymethylation on cytosine would greatly impact the binding of transcription factors (TFs). However, there is a lack of motif finding algorithms with the function to search for motifs with PTMs. In this study, we expend on our previous motif finding pipeline Epigram to provide systematic de novo motif discovery and performance evaluation on methylated DNA motifs. Using the tool, we were able to identified methylated motifs in Arabidopsis DAP-seq data that were previously demonstrated to contain such motifs1. When applied to TF ChIP-seq and DNA methylome data in H1 and GM12878, our method successfully identified novel methylated motifs that can be recognized by the TFs or their co-factors. We also observed spacing constraint between the canonical motif of the TF and the newly discovered methylated motifs, which suggests operative recognition of these cis-elements by collaborative proteins.

Bioinformatics

Systematic identification of cooperation between DNA binding proteins in 3D space

Cooperation between DNA-binding proteins (DBPs) such as transcription factors and chromatin remodeling enzymes plays a pivotal role in regulating gene expression and other biological processes. Such cooperation is often via interaction between DBPs that bind to loci located distal in the linear genome but close in the 3D space, referred as trans-cooperation. Due to the lack of 3D chromosomal structure, identification of DBP cooperation has been limited to those binding to neighbor regions in the linear genome, referred as cis-cooperation. Here we present the first study that integrates protein ChIP-seq and Hi-C data to systematically identify both cis- and trans-cooperation between DBPs. We developed a new network model that allows identification of cooperation between multiple DBPs and reveals cell type specific or independent regulations. Particularly interesting, we have retrieved many known and previously unknown trans-cooperation between DBPs in the chromosomal loops that may be a key factor for influencing 3D chromosomal structure. The software is available at http://wanglab.ucsd.edu/star/DBPnet/index.html.

Bioinformatics

Targeted mutations on 3D hub loci alter spatial interaction environment

Many disease-related genotype variations (GVs) reside in non-gene coding regions and the mechanisms of their association with diseases are largely unknown. A possible impact of GVs on disease formation is to alter the spatial organization of chromosome. However, the relationship between GVs and 3D genome structure has not been studied at the chromosome scale. The kilobase resolution of chromosomal structures measured by Hi-C have provided an unprecedented opportunity to tackle this problem. Here we proposed a network-based method to capture global properties of the chromosomal structure. We uncovered that genome organization is scale free and the genomic loci interacting with many other loci in space, termed as hubs, are critical for stabilizing local chromosomal structure. Importantly, we found that cancer-specific GVs target hubs to drastically alter the local chromosomal interactions. These analyses revealed the general principles of 3D genome organization and provided a new direction to pinpoint genotype variations in non-coding regions that are critical for disease formation.

Biophysics

Full-genome evolutionary histories of selfing, splitting and selection in Caenorhabditis

The nematode Caenorhabditis briggsae is a model for comparative developmental evolution with C. elegans. Worldwide collections of C. briggsae have implicated an intriguing history of divergence among genetic groups separated by latitude, or by restricted geography, that is being exploited to dissect the genetic basis to adaptive evolution and reproductive incompatibility. And yet, the genomic scope and timing of population divergence is unclear. We performed high-coverage whole-genome sequencing of 37 wild isolates of the nematode C. briggsae and applied a pairwise sequentially Markovian coalescent (PSMC) model to 703 combinations of genomic haplotypes to draw inferences about population history, the genomic scope of natural selection, and to compare with 40 wild isolates of C. elegans. We estimate that a diaspora of at least 6 distinct C. briggsae lineages separated from one another approximately 200 thousand generations ago, including the Temperate and Tropical phylogeographic groups that dominate most samples from around the world. Moreover, an ancient population split in its history 2 million generations ago, coupled with only rare gene flow among lineage groups, validates this system as a model for incipient speciation. Low versus high recombination regions of the genome give distinct signatures of population size change through time, indicative of widespread effects of selection on highly linked portions of the genome owing to extreme inbreeding by self-fertilization. Analysis of functional mutations indicates that genomic context, owing to selection that acts on long linkage blocks, is a more important driver of population variation than are the functional attributes of the individually encoded genes.

Genomics

Hybrid origins and the earliest stages of diploidization in the highly successful recent polyploid Capsella bursa-pastoris

Whole genome duplication events have occurred repeatedly during flowering plant evolution, and there is growing evidence for predictable patterns of gene retention and loss following polyploidization. Despite these important insights, the rate and processes governing the earliest stages of diploidization remain poorly understood, and the relative importance of genetic drift, positive selection and relaxed purifying selection in the process of gene degeneration and loss is unclear. Here, we conduct whole genome resequencing in Capsella bursa-pastoris, a recently formed tetraploid with one of the most widespread species distributions of any angiosperm. Whole genome data provide strong support for recent hybrid origins of the tetraploid species within the last 100-300,000 years from two diploid progenitors in the Capsella genus. Major-effect inactivating mutations are frequent, but many were inherited from the parental species and show no evidence of being fixed by positive selection. Despite a lack of large-scale gene loss, we observe a decrease in the efficacy of natural selection genome-wide, due to the combined effects of demography, selfing and genome redundancy from whole genome duplication. Our results suggest that the earliest stages of diploidization are associated with quantitative genome-wide decreases in the strength and efficacy of selection rather than rapid gene loss, and that non-functionalization can receive a 'head start' through a legacy of deleterious variants and differential expression originating in parental diploid populations.

Evolutionary Biology

The Effectiveness of China’s National Forest Protection Program and National-level Nature Reserves, 2000 to 2010: PREPRINT

There is profound interest in knowing the degree to which Chinas institutions are capable of protecting its natural forests and biodiversity in the face of economic and political change. Chinas two most important forest protection policies are its National Forest Protection Program (NFPP) and its National-level Nature Reserves (NNRs). The NFPP was implemented in 17 provinces starting in the year 2000 in response to deforestation-caused flooding. We used MODIS data (MOD13Q1) to estimate forest cover and forest loss across mainland China, and we report that 1.765 million km2 or 18.7% of mainland China was covered in forest (12.3%, canopy cover > 70%) and woodland (6.4%, 40% [&le;] canopy cover < 70%) in 2000. By 2010, a total of 480,203 km2 of forest + woodland was lost, amounting to an annual deforestation rate of 2.7%. The forest-only loss was 127,473 km2, or 1.05% annually. The three most rapidly deforested provinces were outside NFPP jurisdiction, in the southeast. Within the NFPP provinces, the annual forest + woodland loss rate was 2.26%, and the forest-only rate was 0.62%. Because these loss rates are likely overestimates, China appears to have achieved, and even exceeded, its NFPP target of reducing deforestation to 1.1% annually in the target provinces. We also assemble the first-ever polygon dataset for Chinas forested NNRs (n = 237), which covered 74,030 km2 in 2000. Conventional unmatched and covariate-matching analyses both find that about two-thirds of Chinas NNRs exhibit effectiveness in protecting forest cover and that within-NNR deforestation rates are higher in provinces that have higher overall deforestation.

Ecology