Search bioRxivSearch

Biology subjects

Xie, Z.

Publications and source records attributed to Xie, Z..

10 recordsLinked to original sources

Elysium: RNA-seq Alignment in the Cloud

MotivationRNA-sequencing (RNA-seq) is currently the leading technology for genome-wide transcript quantification. Mapping the raw reads to transcript and gene level counts can be achieved by a variety of aligners and pipelines. The diversity of processing options reduces interoperability. In addition, the alignment step requires significant computational resources and basic programming knowledge. Elysium enables users of all skill levels to perform a uniform and free RNA-seq alignment in the cloud.\n\nResultsThe Elysium infrastructure is comprised of four components: A file upload API that enables storage of FASTQ files on Amazon S3 without Amazon credentials; an API to handle the cloud alignment job scheduling for uploaded files; and a graphical user interface (GUI) to provide intuitive access to users that do not have command-line access skills.\n\nAvailabilityThe Elysium source code is available under the Apache Licence 2.0 on GitHub at: https://github.com/maayanlab/elysium\n\nThe service of cloud based RNA-seq alignment is freely accessible through the Elysium GUI at: http://elysium.cloud

bioinformatics

Genetically modified pigs are protected from classical swine fever virus

Classical swine fever (CSF) caused by classical swine fever virus (CSFV) is among the most detrimental diseases, and leads to significant economic losses in the swine industry. Despite efforts by many government authorities try to stamp out the disease from national pig populations, the disease remains widespread. Here, antiviral small hairpin RNAs (shRNAs) were selected and then inserted at the porcine ROSA26 (pROSA26) locus via a CRISPR/Cas9-mediated knock-in strategy. Finally, anti-CSFV transgenic (TG) pigs were produced by somatic nuclear transfer (SCNT). Importantly, in vitro and in vivo viral challenge assays demonstrated that these TG pigs could effectively limit the growth of CSFV and reduced CSFV-associated clinical signs and mortality, and the disease resistance was stably transmitted to F1-generation. The use of these TG pigs can improve the well-being of livestock and substantially reduce virus-related economic losses. Additionally, this antiviral approach may provide a reference for future antiviral research.\n\nAuthor summaryClassical swine fever (CSF), caused by classical swine fever virus (CSFV), and is a highly contagious, often fatal porcine disease with significant economic losses. Due to its economic importance to the pig industry, the biology and pathogenesis of CSFV have been investigated extensively. Despite efforts by many government authorities to stamp out the disease from national pig populations, the disease remains widespread in some regions and seems to be waiting for the reintroduction and the next round of disease outbreaks. These highlight the necessity and urgency of developing more effective approaches to eradicate the challenging CSFV. In this study, we successfully produced anti-CSFV transgenic pigs and confirmed that these transgenic pigs could effectively limit the growth of CSFV in vivo and in vitro and that the disease resistance traits in the TG founders can be stably transmitted to their F1-generation offspring. This study suggests that these TG pigs can improve the well-being of livestock and contribute to offer potential benefits over commercial vaccination. The use of these TG pigs can improve the well-being of livestock and substantially reduce CSFV-related economic losses.

genomics

Improved sgRNA design in bacteria via genome-wide activity profiling

CRISPR/Cas9 is a promising tool in prokaryotic genome engineering, but its success is limited by the widely varying on-target activity of single guide RNAs (sgRNAs). Based on the association of CRISPR/Cas9-induced DNA cleavage with cellular lethality, we systematically profiled sgRNA activity by co-expressing a genome-scale library (~70,000 sgRNAs) with Cas9 or its specificity-improved mutant in E. coli. Based on this large-scale dataset, we constructed a comprehensive and high-density sgRNA activity map, which enables selecting highly active sgRNAs for any locus across the genome in this model organism. We also identified resistant genomic loci with respect to CRISPR/Cas9 activity, notwithstanding the highly accessible DNA in bacterial cells. Moreover, we found that previous sgRNA activity prediction models that were trained on mammalian cell datasets were inadequate when coping with our results, highlighting the key limitations and biases of previous models. We hence developed an integrated algorithm to accurately predict highly effective sgRNAs, aiming to facilitate the design of CRISPR/Cas9-based genome engineering or screenings in bacteria. We also isolated the important sgRNA features that contribute to DNA cleavage and characterized their key differences among wild type Cas9 and its mutant, shedding light on the biophysical mechanisms of the CRISPR/Cas9 system.

microbiology

Regulation by competition: a hidden layer of gene regulatory network

Molecular competition is ubiquitous, essential and multifunctional throughout diverse biological processes. Competition brings about trade-offs of shared limited resources among the cellular components, and it thus introduce a hidden layer of regulatory mechanism by connecting components even without direct physical interactions. By abstracting the analogous competition mechanism behind diverse molecular systems, we built a unified coarse-grained competition motif model to systematically compare experimental evidences in these processes and analyzed general properties shared behind them. We could predict in what molecular environments competition would reveal threshold behavior or display a negative linear dependence. We quantified how competition can shape regulator-target dose-response curve, modulate dynamic response speed, control target expression noise, and introduce correlated fluctuations between targets. This work uncovered the complexity and generality of molecular competition effect, which might act as a hidden regulatory mechanism with multiple functions throughout biological networks in both natural and synthetic systems.

systems biology

Stimulation of the final cell cycle in the stomatal lineage by the cyclin CYCD7;1 under regulation of the MYB transcription factor FOUR-LIPS

Abstract (180 words)Stomatal guard cells are formed through a sequence of asymmetric and symmetric divisions in the epidermis of the sporophyte of most land plants. We show that several D-type cyclins are consecutively activated in the stomatal linage in the epidermis of Arabidopsis thaliana. Whereas CYCD2;1 and CYCD3;2 are activated in the meristemoids early in the lineage, CYCD7;1 is activated before the final division. CYCD7;1 expression peaks in the guard mother cell, where its transcription is modulated by the FOUR-LIPS/MYB88 transcription factor. FOUR-LIPS/MYB88 interacts with the CYCD7;1 promoter and represses CYCD7;1 transcription. CYCD7;1 stimulates the final symmetric division in the stomatal lineage, since guard cell formation is delayed in the cycd7;1 mutant epidermis and guard mother cell (GMC) divisions in four-lips mutant guard mother cells are limited by loss of function of CYCD7;1. Hence, the precise activation of a specific D-type cyclin, CYCD7;1, is required for correct timing of the last symmetric division that creates the stomatal guards cells, and CYCD7;1 expression is regulated by the FLP/MYB pathway that ensures cell cycle arrest in the stomatal guard cells.\n\nSummary StatementThe formation of paired guard cells in the epidermis of the Arabidopsis thaliana shoot, requires the activity of the D-type cyclin CYCD7;1 for the normal timing of the final division.

plant biology

Pooled CRISPR interference screens enable high-throughput functional genomics study and elucidate new rules for guide RNA library design in Escherichia coli

Clustered regularly interspaced short palindromic repeat (CRISPR)/Cas9 technology provides potential advantages in high-throughput functional genomics analysis in prokaryotes over previously established platforms based on recombineering or transposon mutagenesis. In this work, as a proof-of-concept to adopt CRISPR/Cas9 method as a pooled functional genomics analysis platform in prokaryotes, we developed a CRISPR interference (CRISPRi) library consisting of 3,148 single guide RNAs (sgRNAs) targeting the open reading frame (ORF) of 67 genes with known knockout phenotypes and performed pooled screens under two stressed conditions (minimal and acidic medium) in Escherichia coli. Our approach confirmed most of previously described gene-phenotype associations while maintaining < 5% false positive rate, suggesting that CRISPRi screen is both sensitive and specific. Our data also supported the ability of this method to narrow down the candidate gene pool when studying operons, a unique structure in prokaryotic genome. Meanwhile, assessment of multiple loci across treatments enables us to extract several guidelines for sgRNA design for such pooled functional genomics screen. For instance, sgRNAs locating at the first 5% upstream region within ORF exhibit enhanced activity and 10 sgRNAs per gene is suggested to be enough for robust identification of gene-phenotype associations. We also optimized the hit-gene calling algorithm to identify target genes more robustly with even fewer sgRNAs. This work showed that CRISPRi could be adopted as a powerful functional genomics analysis tool in prokaryotes and provided the first guideline for the construction of sgRNA libraries in such applications.\n\nImportanceTo fully exploit the valuable resource of explosive sequenced microbial genomes, high-throughput experimental platform is needed to associate genes and phenotypes at the genome level, giving microbiologists the insight about the genetic structure and physiology of a microorganism. In this work, we adopted CRISPR interference method as a pooled high-throughput functional genomics platform in prokaryotes with Escherichia coli as the model organism. Our data suggested that this method was highly sensitive and specific to map genes with previously known phenotypes, potent to act as a new strategy for high-throughput microbial genetics study with advantages over previously established methods. We also provided the first guideline for the sgRNA library design by comprehensive analysis of the screen data. The concept, gRNA library design rules and open-source scripts of this work should benefit prokaryotic genetics community to apply high-throughput mapping of defined gene set with phenotypes in a broad spectrum of microorganisms.

microbiology

An inducible CRISPR-ON system for controllable gene activation in human pluripotent stem cells

Human pluripotent stem cells (hPSCs) are an important system to study early human development, model human diseases, and develop cell replacement therapies. However, genetic manipulation of hPSCs is challenging and a method to simultaneously activate multiple genomic sites in a controllable manner is sorely needed. Here, we constructed a CRISPR-ON system to efficiently upregulate endogenous genes in hPSCs. A doxycycline (Dox) inducible dCas9-VP64-p65-Rta (dCas9-VPR) transcription activator and a reverse Tet transactivator (rtTA) expression cassette were knocked into the two alleles of the AAVS1 locus to generate an iVPR hESC line. We showed that the dCas9-VPR level could be precisely and reversibly controlled by addition and withdrawal of Dox. Upon transfection of multiplexed gRNA plasmid targeting the NANOG promoter and Dox induction, we were able to control NANOG gene expression from its endogenous locus. Interestingly, an elevated NANOG level did not only promote naive pluripotent gene expression but also enhanced cell survival and clonogenicity, and it enabled integration of hESCs with the inner cell mass (ICM) of mouse blastocysts in vitro. Thus, iVPR cells provide a convenient platform for gene function studies as well as high-throughput screens in hPSCs.

cell biology

MECAT: an ultra-fast mapping, error correction and de novo assembly tool for single-molecule sequencing reads

The high computational cost of current assembly methods for the long, noisy single molecular sequencing (SMS) reads has prevented them from assembling large genomes. We introduce an ultra-fast alignment method based on a novel global alignment score. For large human SMS data, our method is 7X faster than MHAP for pairwise alignment and 15X faster than BLASR for reference mapping. We develop a Mapping, Error Correction and de novo Assembly Tool (MECAT) by integrating our new alignment and error correction methods, with the Celera Assembler. MECAT is capable of producing high quality de novo assembly of large genome from SMS reads with low computational cost. MECAT produces reference-quality assemblies of Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster and reconstructs the human CHM1 genome with 15% longer NG50 in only 7600 CPU core hours using 54X SMS reads and a Chinese Han genome in 19200 CPU core hours using 102X SMS reads.

bioinformatics

HICL table can manipulate all proteins in human complete proteome

BackgroundThe data of human complete proteome in the databases of Universal Protein Resource (UniProt) or National Center for Biotechnology Information(NCBI) were disorderly organized and hardly handled by an ordinary biologist.\n\nResultsThe HICL table enable an ordinary biologist efficiently to handle the human complete proteome with 67911 entries, to get an overview on the distribution of the physicochemical features of all proteins in the human complete proteome, to perceive the details of the distribution patterns of the physicochemical features in some protein family members and protein variants, to find some particular proteins.\n\nMoreover, two discoveries were made via the HICL table: (1) The amino aicds(Asp,Glu) have symmetrical trend of the distributions versus pI, but the amino aicds(Arg, Lys) have local asymmetrical trend of the distributions versus pI in human complete proteome. (2) Protein sequence, besides amino acid properties, can in theory influence the modal distribution of protein isoelectric points.\n\nO_TBL View this table:\norg.highwire.dtl.DTLVardef@89820eorg.highwire.dtl.DTLVardef@1b99fe1org.highwire.dtl.DTLVardef@1af7e13org.highwire.dtl.DTLVardef@7e22acorg.highwire.dtl.DTLVardef@1164e8f_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOTable1:C_FLOATNO O_TABLECAPTIONThe values and ranges of the grouping criteria for every feature\n\nC_TABLECAPTION C_TBL ConclusionI has created the HICL table as a robust tool for orderly managing 67911 proteins in human complete proteome by their physicochemical features, the names and sequences. Any proteins with the particular physicochemical features can be screened out from the human complete proteome via the HICL table. In addition, the unbalanced distribution of the amino aicds(Arg, Lys) in high pI proteins of human complete proteome and the effect of protein sequence on modal distribution of protein isoelectric points have been discovered through the HICL table.

bioinformatics

The unbalanced distribution of the amino aicds(Arg, Lys) in high pI proteins of complete proteomes, its origin and evolution

No evolutionary signature has been found in the complete proteomes of eukaryotes by now, although amino acid composition signature of the complete proteomes was discovered as molecular signature of the adaptation for thermophiles and halophiles. Arginine and lysine respectively have the guanidinium group and the [isin]-amino group as the ionizable side chain groups with different pKa values of about 12.5 and 10.5. The trends of their distribution seem similar in the range of about pIs < 10.0 and diverge in the range of about pIs[&ge;]10.0 in most complete proteomes of 387 species from the three domains of life. The complete proteome of Reticulomyxa filose is the one of only in 287 eukaryotic complete proteomes that has a predominance of the trend of lysine over that of arginine in high pI proteins. The unbalanced distribution of the amino aicds(Arg, Lys) in high pI proteins of complete proteomes may originally come from different pKa values of arginine and lysine and be developed by the influences of an average lysine level of a proteome and evolution. Because of this unbalanced distribution, the pattern of arginine and lysine distribution in high pI proteins of some complete proteomes can form a particular proteomic structure as evolutionary signature that may be shaped by massive natural selection in molecular level from hundreds to ten thousands of proteins in the complete proteomes of many animals and green plants.

bioinformatics