Search bioRxivSearch

Biology subjects

Gao, Z.

Publications and source records attributed to Gao, Z..

12 recordsLinked to original sources

Landscape of stimulation-responsive chromatin across diverse human immune cells

The immune system is controlled by a balanced interplay among specialized cell types transitioning between resting and stimulated states. Despite its importance, the regulatory landscape of this system has not yet been fully characterized. To address this gap, we collected ATAC-seq and RNA-seq data under resting and stimulated conditions for 25 immune cell types from peripheral blood of four healthy individuals, and seven cell types from three fetal thymus samples. We found that stimulation caused widespread chromatin remodeling, including a large class of response elements shared between stimulated B and T cells. Furthermore, several autoimmune traits showed significant heritability in stimulation-responsive elements from distinct cell types, highlighting the critical importance of these cell states in autoimmunity. Use of allele-specific read-mapping identified thousands of variants that alter chromatin accessibility in particular conditions. Notably, variants associated with changes in stimulation-specific chromatin accessibility were not enriched for associations with gene expression regulation in whole blood - a tissue commonly used in eQTL studies. Thus, large-scale maps of variants associated with gene regulation lack a condition important for understanding autoimmunity. As a proof-of-principle we identified variant rs6927172, which links stimulated T cell-specific chromatin dysregulation in the TNFAIP3 locus to ulcerative colitis and rheumatoid arthritis. Overall, our results provide a broad resource of chromatin landscape dynamics and highlight the need for large-scale characterization of effects of genetic variation in stimulated cells.

genomics

Integrated modeling of peptide digestion and detection for the prediction of proteotypic peptides in targeted proteomics

MotivationThe selection of proteotypic peptides, i.e., detectable unique representatives of proteins of interest, is a key step in targeted shotgun proteomics. To date, much effort has been made to predict proteotypic peptides in the absence of mass spectrometry data. However, the performance of existing tools is still unsatisfactory. One crucial reason is their neglect of the close relationship between protein proteolytic digestion and peptide detection.\n\nResultsWe present an algorithm (named AP3) that firstly considers peptide digestion probability as a feature for proteotypic peptide prediction and demonstrated peptide digestion probability is the most important feature for accurate prediction of proteotypic peptides. AP3 showed higher accuracy than existing tools and accurately predicted the proteotypic peptides for a targeted proteomics assay, showing its great potential for assisting the design of targeted proteomics experiments.\n\nAvailability and ImplementationFreely available at http://fugroup.amss.ac.cn/software/AP3/AP3.html.\n\nContactyfu@amss.ac.cn or zhuyunping@gmail.com\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics

Single-cell RNA-seq reveals dynamic transcriptome profiling in human early neural differentiation

BackgroundInvestigating cell fate decision and subpopulation specification in the context of the neural lineage is fundamental to understanding neurogenesis and neurodegenerative diseases. The differentiation process of neural-tube-like rosettes in vitro is representative of neural tube structures, which are composed of radially organized, columnar epithelial cells and give rise to functional neural cells. However, the underlying regulatory network of cell fate commitment during early neural differentiation remains elusive.\n\nResultsIn this study, we investigated the genome-wide transcriptome profile of single cells from six consecutive reprogramming and neural differentiation time points and identified cellular subpopulations present at each differentiation stage. Based on the inferred reconstructed trajectory and the characteristics of subpopulations contributing the most towards commitment to the central nervous system (CNS) lineage at each stage during differentiation, we identified putative novel transcription factors in regulating neural differentiation. In addition, we dissected the dynamics of chromatin accessibility at the neural differentiation stages and revealed active c/s-regulatory elements for transcription factors known to have a key role in neural differentiation as well as for those that we suggest are also involved. Further, communication network analysis demonstrated that cellular interactions most frequently occurred among embryoid body (EB) stage and each cell subpopulation possessed a distinctive spectrum of ligands and receptors associated with neural differentiation which could reflect the identity of each subpopulation.\n\nConclusionsOur study provides a comprehensive and integrative study of the transcriptomics and epigenetics of human early neural differentiation, which paves the way for a deeper understanding of the regulatory mechanisms driving the differentiation of the neural lineage.

developmental biology

Measurement of selective constraint on human gene expression

Gene expression variation is a major contributor to phenotypic variation in human complex traits. Selection on complex traits may therefore be reflected in constraint on gene expression levels. Here, we explore the effects of stabilizing selection on cis-regulatory genetic variation in humans. We analyze patterns of expression variation at copy number variants and find evidence for selection against large increases in gene expression. Using allele-specific expression (ASE) data, we further show evidence of selection against smaller-effect variants. We estimate that, across all genes, singletons in a sample of 122 individuals have approximately 2.5 x greater effects on expression variance than common variants. Despite their increased effect sizes relative to common variants, we estimate that singletons in the sample studied explain, on average, only 5% of the heritability of gene expression from cis-regulatory variants. Finally, we show that genes depleted for loss-of-function variants are also depleted for cis-eQTLs and have low levels of allelic imbalance, confirming tighter constraint on the expression levels of these genes. We conclude that constraint on gene expression is present, but has relatively weak effects on most cis-regulatory variants, thus permitting high levels of gene-regulatory genetic variation.

genomics

LFAQ: towards unbiased label-free absolute protein quantification by predicting peptide quantitative factors

Mass spectrometry (MS) has become a prominent choice for large-scale absolute protein quantification, but its quantification accuracy still has substantial room for improvement. A crucial issue is the bias between the peptide MS intensity and the actual peptide abundance, i.e., the fact that peptides with equal abundance may have different MS intensities. This bias is mainly caused by the diverse physicochemical properties of peptides. Here, we propose a novel algorithm for label-free absolute protein quantification, LFAQ, which can correct the biased MS intensities by using the predicted peptide quantitative factors for all identified peptides. When validated on datasets produced by different MS instruments and data acquisition modes, LFAQ presented accuracy and precision superior to those of existing methods. In particular, it reduced the quantification error by an average of 46% for low-abundance proteins.

bioinformatics

Overlooked roles of DNA damage and maternal age in generating human germline mutations

Although the textbook view is that most germline mutations arise from replication errors, when analyzing large de novo mutation datasets in humans, we find multiple lines of evidence that call that understanding into question. Notably, despite the drastic increase in the ratio of male to female germ cell divisions after the onset of spermatogenesis, even young fathers contribute three times more mutations than young mothers, and this ratio barely increases with parental ages. This surprising finding points to a substantial contribution of damage-induced mutations. Indeed, C to G transversions and CpG transitions, which together constitute one third of all mutations, show genomic distributions and sex-specific age dependencies indicative of doublestrand break repair and methylation-associated damage, respectively. Moreover, the data indicate that maternal age at conception influences the mutation rate both because of the accumulation of damage in oocytes and potentially through an influence on the number of postzygotic mutations.

genetics

Deconvolution of single-cell multi-omics layers reveals regulatory heterogeneity

Integrative analysis of multi-omics layers at single cell level is critical for accurate dissection of cell-to-cell variation within certain cell populations. Here we report scCAT-seq, a technique for simultaneously assaying chromatin accessibility and the transcriptome within the same single cell. We show that the combined single cell signatures enable accurate construction of regulatory relationships between cis-regulatory elements and the target genes at single-cell resolution, providing a new dimension of features that helps direct discovery of regulatory patterns specific to distinct cell identities. Moreover, we generated the first single cell integrated maps of chromatin accessibility and transcriptome in human pre-implantation embryos and demonstrated the robustness of scCAT-seq in the precise dissection of master transcription factors in cells of distinct states during embryo development. The ability to obtain these two layers of omics data will help provide more accurate definitions of \"single cell state\" and enable the deconvolution of regulatory heterogeneity from complex cell populations.

genomics

Sex- and Context-dependent Effects of Oxytocin on Social Reward Processing

We interact socially and form bonds with others because such experiences are rewarding. However, an insecure attachment style or social anxiety can reduce these rewarding effects. The neuropeptide oxytocin (OXT) may facilitate social interactions either by increasing their rewarding experience or by attenuating anxiety, although effects can be sex- and attachment-style dependent. In this study, 64 pairs of same-sex friends completed a social sharing paradigm in a double-blind, placebo-controlled, between-subject design with one friend inside an MRI scanner and the other in a remote behavioral testing room. In this way we could examine whether intranasal-OXT differentially modulated the emotional impact of social sharing and associated neural processing. Additionally, we investigated if OXT effects were modulated by sex and attachment style. Results showed that in women, but not men, OXT increased ratings for sharing stimuli with their friend but not with a stranger, particularly in the friend in the scanner. Corresponding neuroimaging results showed that OXT decreased both amygdala and insula activity as well as their functional connectivity in women when they shared with friends but had the opposite effect in men. On the other hand, OXT did not enhance responses in brain reward circuitry. In the PLC treated group amygdala responses in women when they shared pictures with their friend were positively associated with attachment anxiety and OXT uncoupled this. Our findings demonstrate that OXT facilitates the impact of sharing positive experiences with others in women, but not men, and that this is associated with differential effects on the amygdala and insula and their functional connections. Furthermore, OXT particularly reduced increased amygdala responses during sharing in individuals with higher attachment anxiety. Thus, OXT effects in this context may be due more to reduced anxiety when sharing with a friend than to enhanced social reward.

neuroscience

HDL-AuNPs-BMS nanoparticle conjugates as molecularly targeted therapy for leukemia

In previous work, gold nanoparticles (AuNPs) with adsorbed high-density lipoprotein (HDL) nanoparticles have been utilized to deliver oligonucleotides, yet HDL-AuNPs functionalized with small molecule inhibitors have not been systematically explored. Here, we report an AuNP-based therapeutic system (HDL-AuNPs-BMS) for acute myeloid leukemia (AML) by delivering BMS309403 (BMS), a small molecule that selectively inhibits AML-promoting factor fatty acid binding protein 4 (FABP4). HDL-AuNPs-BMS are synthesized using a gold nanoparticle as template to control conjugate size and ensure a spherical shape to engineer HDL-like nanoparticle containing BMS. The zeta potential and size of the HDL-AuNPs obtained from transmission electron microscopy (TEM) show that the nanoparticles are electrostatically stable and 25 nm in diameter. Functionally, compared to free drug, HDL-AuNPs-BMS conjugates are more readily internalized by AML cells and have more pronounced effect on downregulation of DNA methyltransferase 1 (DNMT1), reduction of global DNA methylation, and restoration of epigenetically-silenced tumor suppressor p15INK4b coupled with AML growth arrest. Importantly, systemic administration of HDL-AuNPs-BMS conjugates into AML-bearing mice inhibits DNMT1-dependent DNA methylation, induces AML cell differentiation and diminishes AML disease progression without obvious side effects. In summary, these data, for the first time, demonstrate HDL-AuNPs as an effective delivery platform with great potential to attach distinct inhibitors, and HDL-AuNPs-BMS conjugates as a promising therapeutic platform to treat leukemia.

cancer biology

Frequent Non-Allelic Gene Conversion On The Human Lineage And Its Effect On The Divergence Of Gene Duplicates

Gene conversion is the copying of genetic sequence from a \"donor\" region to an \"acceptor\". In non-allelic gene conversion (NAGC), the donor and the acceptor are at distinct genetic loci. Despite the role NAGC plays in various genetic diseases and the concerted evolution of gene families, the parameters that govern NAGC are not well-characterized. Here, we survey duplicate gene families and identify converted tracts in 46% of them. These conversions reflect a large GC-bias of NAGC. We develop a sequence evolution model that leverages substantially more information in duplicate sequences than used by previous methods and use it to estimate the parameters that govern NAGC in humans: a mean converted tract length of 250bp and a probability of 2.5x10-7 per generation for a nucleotide to be converted (an order of magnitude higher than the point mutation rate). Despite this high baseline rate, we show that NAGC slows down as duplicate sequences diverge--until an eventual \"escape\" of the sequences from its influence. As a result, NAGC has a small average effect on the sequence divergence of duplicates. This work improves our understanding of the NAGC mechanism and the role that it plays in the evolution of gene duplicates.

genomics

Worldwide Population Structure Of Escherichia coli Reveals Two Major Subspecies

Recombination is one of the most important mechanisms of prokaryotic species evolution but its exact roles are still in debate. Here we try to infer genome-wide recombination events within a species uti-lizing a dataset of 104 complete genomes of Escherichia coli from diverse origins, among which 45 from world-wide animal-hosts are in-house sequenced using SMRT (single-molecular real time) technology.Two major clades are identified based on evidences of ecological and physiological characteristics, as well as distinct genomic features implying scarce inter-clade genetic exchange. By comparing the synteny of identical fragments genome-widely searched for each genome pair, we achieve a fine-scale map of re-combination within the population. The recombination is rather extensive within clade, which is able to break linkages between genes but does not interrupt core genome framework and primary metabolic port-folios possibly due to natural selection for physiological compatibility and ecological fitness. Meanwhile,the recombination between clades declines drastically as the phylogenetic distance increases, generally 10-fold reduced than those of the intra-clade, which establishes genetic barrier between clades. These empirical data of recombination suggest its critical role in the early stage of speciation, where recombina-tion rate differs according to phylogentic distance. The extensive intra-clade recombination coheres sister strains into a quasi-sexual group and optimizes genes or alleles to streamline physiological activities,whereas shapely declined inter-clade recombination split the population into clades adaptive to divergent ecological niches.\n\nSignificance StatementRoles of recombination in species evolution have been debated for decades due to difficulties in inferring recombination events during the early stage of speciation, especially when recombination is always complicated by frequent gene transfer events of bacterial genomes. Based on 104 high-quality complete E. coli genomes, we infer gene-centric dynamics of recombination in the formation of two E. coli clades or subpopulations, and recombination is found to be rather intensive in a within-clade fashion, which forces them to be quasi-sexual. The recombination events can be mapped among individual genomes in the context of genes and their variations; decreased between-clade and increased intra-claderecombination engender a genetic barrier that further encourages clade-specific secondary metabolic portfolios for better environmental adaptation. Recombination is thus a major force that accelerates bacterial evolution to fit ecological diversity.

microbiology

The population genetics of human disease: the case of recessive, lethal mutations

Do the frequencies of disease mutations in human populations reflect a simple balance between mutation and purifying selection? What other factors shape the prevalence of disease mutations? To begin to answer these questions, we focused on one of the simplest cases: recessive mutations that alone cause lethal diseases or complete sterility. To this end, we generated a hand-curated set of 417 Mendelian mutations in 32 genes, reported to cause a recessive, lethal Mendelian disease. We then considered analytic models of mutation-selection balance in infinite and finite populations of constant sizes and simulations of purifying selection in a more realistic demographic setting, and tested how well these models fit allele frequencies estimated from 33,370 individuals of European ancestry. In doing so, we distinguished between CpG transitions, which occur at a substantially elevated rate, and three other mutation types. The observed frequency for CpG transitions is slightly higher than expectation but close, whereas the frequencies observed for the three other mutation types are an order of magnitude higher than expected. This discrepancy is even larger when subtle fitness effects in heterozygotes or lethal compound heterozygotes are taken into account. In principle, higher than expected frequencies of disease mutations could be due to widespread errors in reporting causal variants, compensation by other mutations, or balancing selection. It is unclear why these factors would have a greater impact on variants with lower mutation rates, however. We argue instead that the unexpectedly high frequency of disease mutations and the relationship to the mutation rate likely reflect an ascertainment bias: of all the mutations that cause recessive lethal diseases, those that by chance have reached higher frequencies are more likely to have been identified and thus to have been included in this study. Beyond the specific application, this study highlights the parameters likely to be important in shaping the frequencies of Mendelian disease alleles.\n\nAuthor SummaryWhat determines the frequencies of disease mutations in human populations? To begin to answer this question, we focus on one of the simplest cases: mutations that cause completely recessive, lethal Mendelian diseases. We first review theory about what to expect from mutation and selection in a population of finite size and further generate predictions based on simulations using a realistic demographic scenario of human evolution. For a highly mutable type of mutations, such as transitions at CpG sites, we find that the predictions are close to the observed frequencies of recessive lethal disease mutations. For less mutable types, however, predictions substantially under-estimate the observed frequency. We discuss possible explanations for the discrepancy and point to a complication that, to our knowledge, is not widely appreciated: that there exists ascertainment bias in disease mutation discovery. Specifically, we suggest that alleles that have been identified to date are likely the ones that by chance have reached higher frequencies and are thus more likely to have been mapped. More generally, our study highlights the factors that influence the frequencies of Mendelian disease alleles.

genetics