Search bioRxivSearch

Biology subjects

Ye, Y.

Publications and source records attributed to Ye, Y..

10 recordsLinked to original sources

An enriched network motif family regulates multistep cell fate transitions with restricted reversibility

Multistep cell fate transitions with stepwise changes of transcriptional profiles are common to many developmental, regenerative and pathological processes. The multiple intermediate cell lineage states can serve as differentiation checkpoints or branching points for channeling cells to more than one lineages. However, mechanisms underlying these transitions remain elusive. Here, we explored gene regulatory circuits that can generate multiple intermediate cellular states with stepwise modulations of transcription factors. With unbiased searching in the network topology space, we found a motif family containing a large set of networks can give rise to four attractors with the stepwise regulations of transcription factors, which limit the reversibility of three consecutive steps of the lineage transition. We found that there is an enrichment of these motifs in a transcriptional network controlling the early T cell development, and a mathematical model based on this network recapitulates multistep transitions in the early T cell lineage commitment. By calculating the energy landscape and minimum action paths for the T cell model, we quantified the stochastic dynamics of the critical factors in response to the differentiation signal with fluctuations. These results are in good agreement with experimental observations and they suggest the stable characteristics of the intermediate states in the T cell differentiation. These dynamical features may help to direct the cells to correct lineages during development. Our findings provide general design principles for multistep cell linage transitions and new insights into the early T cell development. The network motifs containing a large family of topologies can be useful for analyzing diverse biological systems with multistep transitions.\n\nAuthor summaryThe functions of cells are dynamically controlled in many biological processes including development, regeneration and disease progression. Cell fate transition, or the switch of cellular functions, often involves multiple steps. The intermediate stages of the transition provide the biological systems with the opportunities to regulate the transitions in a precise manner. These transitions are controlled by key regulatory genes of which the expression shows stepwise patterns, but how the interactions of these genes can determine the multistep processes were unclear. Here, we present a comprehensive analysis on the design principles of gene circuits that govern multistep cell fate transition. We found a large network family with common structural features that can generate systems with the ability to control three consecutive steps of the transition. We found that this type of networks is enriched in a gene circuit controlling the development of T lymphocyte, a crucial type of immune cells. We performed mathematical modeling using this gene circuit and we recapitulated the stepwise and irreversible loss of stem cell properties of the developing T lymphocytes. Our findings can be useful to analyze a wide range of gene regulatory networks controlling multistep cell fate transitions.

systems biology

The design and evaluation of a Bayesian system for detecting and characterizing outbreaks of influenza

1The prediction and characterization of outbreaks of infectious diseases such as influenza remains an open and important problem. This paper describes a framework for detecting and characterizing outbreaks of influenza and the results of testing it on data from ten outbreaks collected from two locations over five years. We model outbreaks with compartment models and explicitly model non-influenza influenza-like illnesses.

epidemiology

Preliminary Study of the Potential Mechanism of CTSD in the Aging Process of Sepiella japonica: Fundamental Function Analysis

Cathepsin D, a kind of endopeptidase, can degrade peptides and proteins in lysosomes, which are involved in cell apoptosis. Previous transcriptome analysis of optic glands of Sepiella japonica across four growth stages, expression of a cathepsin D-like segment was found to be significantly different. Based on the complete cDNA sequence of S. japonica, the CTSD gene (also called sjCTSD, GenBank accession no. KY745896.1) was cloned using RACE amplification; this gene is 1389 bp in length and encodes proteins composed of 393 amino acids. Spatio-temporal expression profiles of the sjCTSD gene were determined using qPCR assays, which showed that the expression levels of sjCTSD constantly increased across four growth stages in 9 of 11 tissues that were investigated. In the optic glands, as well as pancreas and liver cells, sjCTSD expression levels sharply increased during the post-spawning phase. To investigate the potential role of the sjCTSD gene in aging progress, we constructed the prokaryotic expression vector of pET28a/sjCTSD. After induced by IPTG, the recombinant protein sjCTSD was obtained in the form of inclusion bodies, with a molecular size of approximately 40.3 kDa, the inclusion body of sjCTSD can be converted to soluble protein through the denaturation and renaturation. The results of functional experiments showed that sjCTSD could degrade bovine hemoglobin under acidic conditions, and inhibit the growth of Escherichia coli and Vibrio alginolyticus, which was speculated that the increased expression of sjCTSD may help inhibit the invasion of pathogenic bacteria with the immune function of cuttlefish declines during the aging process. To a certain extent, these results indicated the potential functional role of the sjCTSD gene in the aging process of S. japonica. This study provides insights to further understand the roles of lysosomal proteins on anti-aging effects in S. japonica and other cephalopoda species.

genetics

AICM: A Genuine Framework for Correcting Inconsistency Between Large Pharmacogenomics Datasets

The inconsistency of open pharmacogenomics datasets produced by different studies limits the usage of pharmacogenomics in biomarker discovery. Investigation of multiple pharmacogenomics datasets confirmed that the pairwise sensitivity data correlation between drugs, or rows, across different studies (drug-wise) is relatively low, while the pairwise sensitivity data correlation between cell-lines, or columns, across different studies (cell-wise) is considerably strong. This common interesting observation across multiple pharmacogenomics datasets suggests the existence of subtle consistency among the different studies (i.e., strong cell-wise correlation). However, significant noises are also shown (i.e., weak drug-wise correlation) and have prevented researchers from comfortably using the data directly. Motivated by this observation, we propose a novel framework for addressing the inconsistency between large-scale pharmacogenomics data sets. Our method can significantly boost the drug-wise correlation and can be easily applied to re-summarized and normalized datasets proposed by others. We also investigate our algorithm based on many different criteria to demonstrate that the corrected datasets are not only consistent, but also biologically meaningful. Eventually, we propose to extend our main algorithm into a framework, so that in the future when more data-sets become publicly available, our framework can hopefully offer a \"ground-truth\" guidance for references.

bioinformatics

MSTD: an efficient method for detecting multi-scale topological domains from symmetric and asymmetric 3D genomic maps

The chromosome conformation capture (3C) technique and its variants have been employed to reveal the existence of a hierarchy of structures in three-dimensional (3D) chromosomal architecture, including compartments, topologically associating domains (TADs), sub-TADs and chromatin loops. However, existing methods for domain detection were only designed based on symmetric Hi-C maps, ignoring long-range interaction structures between domains. To this end, we proposed a generic and efficient method to identify multi-scale topological domains (MSTD), including cis- and trans-interacting regions, from a variety of 3D genomic datasets. We first applied MSTD to detect promoter-anchored interaction domains (PADs) from promoter capture Hi-C datasets across 17 primary blood cell types. The boundaries of PADs are significantly enriched with one or the combination of multiple epigenetic factors. Moreover, PADs between functionally similar cell types are significantly conserved in terms of domain regions and expression states. Cell type-specific PADs involve in distinct cell type-specific activities and regulatory events by dynamic interactions within them. We also employed MSTD to define multi-scale domains from typical symmetric Hi-C datasets and illustrated its distinct superiority to the-state-of-art methods in terms of accuracy, flexibility and efficiency.

bioinformatics

CRISPRs for strain tracking and its application to microbiota transplantation data analysis

CRISPR-Cas systems are adaptive immune systems naturally found in bacteria and archaea. Bacteria and archaea use these systems to defend against invaders, including phages, plasmids and other mobile genetic elements. Relying on integration of invader sequences (protospacers) into CRISPR loci (forming spacers flanked by repeats), CRISPR-Cas systems store genetic memory of past invasions. While CRISPR-Cas systems have evolved in response to invading mobile elements, invaders have also developed mechanisms to avoid detection. As a result of arms-race between CRISPR-Cas systems and their targets, the CRISPR arrays typically undergo rapid turnover of the spacers with removal of old spacers and acquisition of new ones. Additionally, different individuals rarely share spacers amongst their microbiome. In this paper, we developed a pipeline (called CRISPRtrack) for strain tracking based on CRISPR spacer content, and applied it to fecal transplantation microbiome data to study the retention of donor strains in recipients. Our results demonstrate the potential use of CRISPRs as a simple yet effective tool for donor strain tracking in fecal transplantation, and also as a general purpose tool for quantifying microbiome similarity.

bioinformatics

Deconvolution of single-cell multi-omics layers reveals regulatory heterogeneity

Integrative analysis of multi-omics layers at single cell level is critical for accurate dissection of cell-to-cell variation within certain cell populations. Here we report scCAT-seq, a technique for simultaneously assaying chromatin accessibility and the transcriptome within the same single cell. We show that the combined single cell signatures enable accurate construction of regulatory relationships between cis-regulatory elements and the target genes at single-cell resolution, providing a new dimension of features that helps direct discovery of regulatory patterns specific to distinct cell identities. Moreover, we generated the first single cell integrated maps of chromatin accessibility and transcriptome in human pre-implantation embryos and demonstrated the robustness of scCAT-seq in the precise dissection of master transcription factors in cells of distinct states during embryo development. The ability to obtain these two layers of omics data will help provide more accurate definitions of \"single cell state\" and enable the deconvolution of regulatory heterogeneity from complex cell populations.

genomics

Construction and analysis of mRNA, miRNA, lncRNA, and TF regulatory networks reveal the key genes in prostate cancer

Purpose: Prostate cancer (PCa) causes a common male urinary system malignant tumour, and the molecular mechanisms of PCa remain poorly understood. This study aims to investigate the underlying molecular mechanisms of PCa with bioinformatics.\n\nMethods: Original gene expression profiles were obtained from the GSE64318 and GSE46602 datasets in the Gene Expression Omnibus (GEO). We conducted differential screens of the expression of genes (DEGs) between two groups using the R software limma package. The interactions between the differentially expressed miRNAs, mRNAs and lncRNAs were predicted and merged with the target genes. Co-expression of the miRNAs, lncRNAs and mRNAs were selected to construct the mRNA-miRNA and-lncRNA interaction networks. Gene Ontology (GO) and Kyoto Encyclopaedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed for the DEGs. The protein-protein interaction (PPI) networks were constructed, and the transcription factors were annotated. The expression of hub genes in the TCGA datasets was verified to improve the reliability of our analysis.\n\nResults: The results demonstrated that 60 miRNAs, 1578 mRNAs and 61 lncRNAs were differentially expressed in PCa. The mRNA-miRNA-lncRNA networks were composed of 5 miRNA nodes, 13 lncRNA nodes, and 45 mRNA nodes. The DEGs were mainly enriched in the nuclei and cytoplasm and were involved in the regulation of transcription, related to sequence-specific DNA binding, and participated in the regulation of the PI3K-Akt signalling pathway. These pathways are related to cancer and focal adhesion signalling pathways. Furthermore, we found that 5 miRNAs, 6 lncRNAs, 6 mRNAs and 2 TFs play important regulatory roles in the interaction network. The expression levels of EGFR, VEGFA, PIK3R1, DLG4, TGFBR1 and KIT were significantly different between PCa and normal prostate tissue.\n\nConclusion: Based on the current study, large-scale effects of interrelated mRNAs, miRNAs, lncRNAs, and TFs were revealed and a model for predicting the mechanism of PCa was provided. This study provides new insight for the exploration of the molecular mechanisms of PCa and valuable clues for further research.

bioinformatics

Distinctive types of postzygotic single-nucleotide mosaicisms in healthy individuals revealed by genome-wide profiling of multiple organs

Postzygotic single-nucleotide mosaicisms (pSNMs) have been extensively studied in tumors and are known to play critical roles in tumorigenesis. However, the patterns and origin of pSNMs in normal organs of healthy humans remain largely unknown. Using whole-genome sequencing and ultra-deep amplicon re-sequencing, we identified and validated 164 pSNMs from 27 postmortem organ samples obtained from five healthy donors. The mutant allele fractions ranged from 1.0% to 29.7%. Inter- and intra-organ comparison revealed two distinctive types of pSNMs, with about half originating during early embryogenesis (embryonic pSNMs) and the remaining more likely to result from clonal expansion events that had occurred more recently (clonal expansion pSNMs). Compared to clonal expansion pSNMs, embryonic pSNMs had higher proportion of C>T mutations with elevated mutation rate at CpG sites. We observed differences in replication timing between these two types of pSNMs, with embryonic and clonal expansion pSNMs enriched in early- and late-replicating regions, respectively. An increased number of embryonic pSNMs were located in open chromatin states and topologically associating domains that transcribed embryonically. Our findings provide new insights into the origin and spatial distribution of postzygotic mosaicism during normal human development.\n\nAuthor SummaryGenomic mosaicism led by postzygotic mutation is the major cause of cancers and many non-cancer developmental disorders. Theoretically, postzygotic mutations should be accumulated during the developmental process of healthy individuals, but the genome-wide characterization of postzygotic mosaicisms across many organ types of the same individual remained limited. In this study, we identified and validated two types of postzygotic mosaicism from the whole-genomes of 27 organs obtained from five healthy donors. We further found that the postzygotic mosaicisms arising during early embryogenesis and later clonal expansion events show distinct genomic patterns in mutation spectrum, replication timing, and chromatin status.

genomics

The structural basis for regulation of the nucleo-cytoplasmic distribution of Bag6 by TRC35

The metazoan protein BCL2-associated athanogene cochaperone 6 (Bag6) forms a hetero-trimeric complex with ubiquitin-like 4A (Ubl4A) and transmembrane domain recognition complex 35 (TRC35). This Bag6 complex is involved in tail-anchored protein targeting and various protein quality control pathways in the cytosol as well as regulating transcription and histone methylation in the nucleus. Here we present a crystal structure of Bag6 and its cytoplasmic retention factor TRC35, revealing that TRC35 is remarkably conserved throughout opisthokont lineage except at the C-terminal Bag6-binding groove, which evolved to accommodate a novel metazoan factor Bag6. Remarkably, while TRC35 and its fungal homolog guided entry of tail-anchored protein 4 (Get4) utilize a conserved hydrophobic patch to bind their respective C-terminal binding partners Bag6 and Get5, Bag6 wraps around TRC35 on the opposite face relative to the Get4-5 interface. We further demonstrate that the residues involved in TRC35 binding are not only critical for occluding the Bag6 nuclear localization sequence from karyopherin binding to retain Bag6 in the cytosol, but also for preventing TRC35 from succumbing to RNF126-mediated ubiquitylation and degradation. The results provide a mechanism for regulation of Bag6 nuclear localization and the functional integrity of the Bag6 complex in the cytosol.

biochemistry