Search bioRxivSearch

Biology subjects

Zeng, J.

Publications and source records attributed to Zeng, J..

At least 19 recordsLinked to original sources

The effect of X-linked dosage compensation on complex trait variation

Quantitative genetics theory predicts that X-chromosome dosage compensation between sexes will have a detectable effect on the amount of genetic and therefore phenotypic trait variances at associated loci in males and females. Here, we systematically examine the role of dosage compensation in complex trait variation in humans in 20 complex traits in a sample of more than 450,000 individuals from the UK Biobank and in 1,600 gene expression traits from a sample of 2,000 individuals as well as across-tissue gene expression from the GTEx resource. We find, on average, twice as much genetic variation for complex traits due to X-linked loci in males compared to females, consistent with a negligible effect of predicted escape from X-inactivation on complex trait variation across traits and also detect biologically relevant X-linked heterogeneity between the sexes for a number of complex traits.

genetics

A genetically-encoded fluorescent sensor enables rapid and specific detection of dopamine in flies, fish, and mice

Dopamine (DA) is a central monoamine neurotransmitter involved in many physiological and pathological processes. A longstanding yet largely unmet goal is to measure DA changes reliably and specifically with high spatiotemporal precision, particularly in animals executing complex behaviors. Here we report the development of novel genetically-encoded GPCR-Activation-Based-DA (GRABDA) sensors that enable these measurements. In response to extracellular DA rises, GRABDA sensors exhibit large fluorescence increases ({Delta}F/F0[~]90%) with sub-second kinetics, nanomolar to sub-micromolar affinities, and excellent molecular specificity. Importantly, GRABDA sensors can resolve a single-electrical-stimulus evoked DA release in mouse brain slices, and detect endogenous DA release in the intact brains of flies, fish, and mice. In freely-behaving mice, GRABDA sensors readily report optogenetically-elicited nigrostriatal DA release and depict dynamic mesoaccumbens DA changes during Pavlovian conditioning or during sexual behaviors. Thus, GRABDA sensors enable spatiotemporal precise measurements of DA dynamics in a variety of model organisms while exhibiting complex behaviors.

neuroscience

Integrating Hi-C and FISH data for modeling 3D organizations of chromosomes

The new advances in various experimental techniques that provide complementary in-formation about the spatial conformations of chromosomes have inspired researchers to develop computational methods to fully exploit the merits of individual data sources and combine them to improve the modeling of chromosome structure. In this paper, we propose GEM-FISH, a first method for reconstructing the 3D models of chromosomes through systematically integrating both Hi-C and FISH data with the prior biophysical knowledge of a polymer model. Comprehensive tests on a set of chromosomes for which both Hi-C and FISH data were available have demonstrated that GEM-FISH can reconstruct the 3D models of chromosomes with more accurate spatial organizations of TADs and compartments than using only Hi-C data. In addition, GEM-FISH can accurately capture the spatial proximity of loop loci and the colocalization of loci from the same sub-compartments. Moreover, our reconstructed 3D models of chromosomes revealed novel patterns of spatial distributions of super-enhancers which can provide useful insights into understanding the functional roles of these super-enhancers in gene regulation. All these results demonstrated that, through integrating both Hi-C and FISH data into a unified framework, GEM-FISH can provide a better tool for modeling the 3D organizations of chromosomes than using the Hi-C data alone.

bioinformatics

A genetically-encoded fluorescent acetylcholine indicator

Acetylcholine (ACh) regulates a diverse array of physiological processes throughout the body, yet cholinergic transmission in the majority of tissues/organs remains poorly understood due primarily to the limitations of available ACh-monitoring techniques. We developed a family of G-protein-coupled receptor activation-based ACh sensors (GACh) with sensitivity, specificity, signal-to-noise ratio, kinetics and photostability suitable for monitoring ACh signals in vitro and in vivo. GACh sensors were validated with transfection, viral and/or transgenic expression in a dozen types of neuronal and non-neuronal cells prepared from several animal species. In all preparations, GACh sensors selectively responded to exogenous and/or endogenous ACh with robust fluorescence signals that were captured by epifluorescent, confocal and/or two-photon microscopy. Moreover, analysis of endogenous ACh release revealed firing pattern-dependent release and restricted volume transmission, resolving two long-standing questions about central cholinergic transmission. Thus, GACh sensors provide a user-friendly, broadly applicable toolbox for monitoring cholinergic transmission underlying diverse biological processes.

neuroscience

Low nutrient levels reduce the fitness cost of MexCD-OprJ efflux pump overexpression in ciprofloxacin-resistant Pseudomonas aeruginosa

The long-term persistence of antibiotic resistance in the environment is a public health concern. Expression of an efflux pump, an important mechanism of resistance to antibiotics, is usually associated with a fitness cost in bacteria. In this study, we aimed to determine why antibiotic resistance conferred by overexpression of an efflux pump persists in environments such as drinking and source water in which antibiotic selective pressure may be very low or even absent. Competition experiments between wild-type Pseudomonas aeruginosa and ciprofloxacin-resistant mutants revealed that the fitness cost of ciprofloxacin resistance (strains cip_1, cip_2, and cip_3) significantly decreased (P < 0.05) under low-nutrient (0.5 mg/l total organic carbon (TOC)) relative to high-nutrient (500 mg/l TOC) conditions. Mechanisms underlying this fitness cost were analyzed. MexD gene expression in resistant bacteria (cip_3 strain) was significantly lower (P < 0.05) in low-nutrient conditions, with 10 mg/l TOC (8.01 {+/-} 0.82-fold), than in high-nutrient conditions, with 500 mg/l TOC (48.89 {+/-} 4.16-fold). Moreover, rpoS gene expression in resistant bacteria (1.36 {+/-} 0.13-fold) was significantly lower (P < 0.05) than that in the wild-type strain (2.78 {+/-} 0.29-fold) under low-nutrient conditions (10 mg/l TOC), suggesting a growth advantage. Furthermore, the difference in metabolic activity between the two competing strains was significantly smaller (P < 0.05) in low-nutrient conditions (5 and 0.5 mg/l TOC). These results suggest that nutrient levels are a key factor in determining the persistence and spread of antibiotic resistance conferred by efflux pumps in the natural environment with trace amounts or no antibiotics.\n\nImportanceThe widespread of antibiotic resistance has led to an increasing concern about the environmental and public health risks. Mechanisms associated with antibiotic resistance including efflux pumps often increase bacterial fitness cost. Our study showed that the fitness cost of ciprofloxacin resistance conferred by overexpression of MexCD-OprJ efflux pump significantly decreased under low-nutrient relative to high-nutrient conditions. The significance of our research is to reveal that nutrient levels are key factor in determining the persistence of antibiotic resistance conferred by efflux pumps under conditions with trace amounts or no antibiotics, which can be mediated by some mechanisms including MexD gene expression, SOURs differences, and rpoS gene regulation.

microbiology

Novel susceptibility loci and genetic regulation mechanisms for type 2 diabetes

We conducted a meta-analysis of genome-wide association studies (GWAS) with [~]16 million genotyped/imputed genetic variants in 62,892 type 2 diabetes (T2D) cases and 596,424 controls of European ancestry. We identified 139 common and 4 rare (minor allele frequency < 0.01) variants associated with T2D, 42 of which (39 common and 3 rare variants) were independent of the known variants. Integration of the gene expression data from blood (n = 14,115 and 2,765) and other T2D-relevant tissues (n = up to 385) with the GWAS results identified 33 putative functional genes for T2D, three of which were targeted by approved drugs. A further integration of DNA methylation (n = 1,980) and epigenomic annotations data highlighted three putative T2D genes (CAMK1D, TP53INP1 and ATP5G1) with plausible regulatory mechanisms whereby a genetic variant exerts an effect on T2D through epigenetic regulation of gene expression. We further found evidence that the T2D-associated loci have been under purifying selection.

genetics

Identifying gene targets for brain-related traits using transcriptomic and methylomic data from blood

Understanding the difference in genetic regulation of gene expression between brain and blood is important for discovering genes associated with brain-related traits and disorders. Here, we estimate the correlation of genetic effects at the top associated cis-expression (cis-eQTLs or cis-mQTLs) between brain and blood for genes expressed (or CpG sites methylated) in both tissues, while accounting for errors in their estimated effects (rb). Using publicly available data (n = 72 to l,366), we find that the genetic effects of cis-eQTLs (PeQTL < 5x10-8) or mQTLs (PmQTL < 1x10-10) are highly correlated between independent brain and blood samples ([Formula] with SE = 0.015 for cis-eQTL and [Formula] with SE = 0.006 for cis-mQTLs). Using meta-analyzed brain eQTL/mQTL data (n = 526 to 1,194), we identify 61 genes and 167 DNA methylation (DNAm) sites associated with 4 brain-related traits and disorders. Most of these associations are a subset of the discoveries (97 genes and 295 DNAm sites) using data from blood with larger sample sizes (n = l,980 to 14,115). We further find that cis-eQTLs with tissue-specific effects are approximately uniformly distributed across all the functional annotation categories, and that mean difference in gene expression level between brain and blood is almost independent of the difference in the corresponding cis-eQTL effect. Our results demonstrate the gain of power in gene discovery for brain-related phenotypes using blood cis-eQTL or cis-mQTL data with large sample sizes.

genetics

TFmapper: A tool for searching putative factors regulating gene expression using ChIP-seq data

BackgroundNext-generation sequencing coupled to chromatin immunoprecipitation (ChIP-seq), DNase I hypersensitivity (DNase-seq) and the transposase-accessible chromatin assay (ATAC-seq) has generated enormous amounts of data, markedly improved our understanding of the transcriptional and epigenetic control of gene expression. To take advantage of the availability of such datasets and provide clues on what factors, including transcription factors, epigenetic regulators and histone modifications, potentially regulates the expression of a gene of interest, a tool for simultaneous queries of multiple datasets using symbols or genomic coordinates as search terms is needed.\n\nResultsIn this study, we annotated the peaks of thousands of ChIP-seq datasets generated by ENCODE project, or ChIP-seq/DNase-seq/ATAC-seq datasets deposited in Gene Expression Omnibus and curated by CistromeDB; We built a MySQL database called TFmapper containing the annotations and associated metadata, allowing users without bioinformatics expertise to search across thousands of datasets to identify factors targeting a genomic region/gene of interest in a specified sample through a web interface. Users can also visualize multiple peaks in genome browsers and download the corresponding sequences.\n\nConclusionTFmapper will help users explore the vast amount of publicly available ChIP-seq/DNase-seq/ATAC-seq data, and perform integrative analyses to understand the regulation of a gene of interest. The web server is freely accessible at http://www.tfmapper.org/.

bioinformatics

NeoDTI: Neural integration of neighborinformation from a heterogeneous network fordiscovering new drug-target interactions

MotivationAccurately predicting drug-target interactions (DTIs) in silico can guide the drug discovery process and thus facilitate drug development. Computational approaches for DTI prediction that adopt the systems biology perspective generally exploit the rationale that the properties of drugs and targets can be characterized by their functional roles in biological networks.\n\nResultsInspired by recent advance of information passing and aggregation techniques that generalize the convolution neural networks (CNNs) to mine large-scale graph data and greatly improve the performance of many network-related prediction tasks, we develop a new nonlinear end-to-end learning model, called NeoDTI, that integrates diverse information from heterogeneous network data and automatically learns topology-preserving representations of drugs and targets to facilitate DTI prediction. The substantial prediction performance improvement over other state-of-the-art DTI prediction methods as well as several novel predicted DTIs with evidence supports from previous studies have demonstrated the superior predictive power of NeoDTI. In addition, NeoDTI is robust against a wide range of choices of hyperparameters and is ready to integrate more drug and target related information (e.g., compound-protein binding affinity data). All these results suggest that NeoDTI can offer a powerful and robust tool for drug development and drug repositioning.\n\nAvailability and implementationThe source code and data used in NeoDTI are available at: https://github.com/FangpingWan/NeoDTI.\n\nContactzengjy321@tsinghua.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

DeepHINT: Understanding HIV-1 integration via deep learning with attention

MotivationHuman immunodeficiency virus type 1 (HIV-1) genome integration is closely related to clinical latency and viral rebound. In addition to human DNA sequences that directly interact with the integration machinery, the selection of HIV integration sites has also been shown to depend on the heterogeneous genomic context around a large region, which greatly hinders the prediction and mechanistic studies of HIV integration.\n\nResultsWe have developed an attention-based deep learning framework, named DeepHINT, to simultaneously provide accurate prediction of HIV integration sites and mechanistic explanations of the detected sites. Extensive tests on a high-density HIV integration site dataset showed that DeepHINT can outperform conventional modeling strategies by automatically learning the genomic context of HIV integration solely from primary DNA sequence information. Systematic analyses on diverse known factors of HIV integration further validated the biological relevance of the prediction result. More importantly, in-depth analyses of the attention values output by DeepHINT revealed intriguing mechanistic implications in the selection of HIV integration sites, including potential roles of several basic helix-loop-helix (bHLH) transcription factors and zinc-finger proteins. These results established DeepHINT as an effective and explainable deep learning framework for the prediction and mechanistic study of HIV integration.\n\nAvailabilityDeepHINT is available as an open-source software and can be downloaded from https://github.com/nonnerdling/DeepHINT\n\nContactlzhang20@mail.tsinghua.edu.cn and zengjy321@tsinghua.edu.cn

bioinformatics

Three classes of response elements for human PRC2 and MLL1/2-trithorax complexes

Polycomb group (PcG) and trithorax group (TrxG) proteins are essential for maintaining epigenetic memory in both embryonic stem cells and differentiated cells. To date, how they are localized to hundreds of specific target genes within a vertebrate genome had remained elusive. Here, by focusing on short cis-acting DNA elements of single functions, we discovered, for the first time, to our knowledge, three classes of response elements in human genome: PcG response elements (PREs), MLL1/2-TrxG response elements (TREs) and PcG/TrxG response elements (P/TREs). We further demonstrated that, in contrast to their proposed roles in recruiting PcG proteins to PREs, YY1 and CpG islands are specifically enriched in TREs and P/TREs, but not in PREs. The three classes of response elements as unraveled in this study open new doors for a deeper understanding of PcG and TrxG mechanisms in vertebrates.

biochemistry

HPCDb: an integrated database of pancreatic cancer

We have established a database of Human Pancreatic Cancer (HPCDb) through effectively mining, extracting, analyzing, and integrating PC-related genes, single-nucleotide polymorphisms (SNPs), and microRNAs (miRNAs), now available online at http://www.pancancer.org/. Data were extracted from established databases, [&ge;]5 published literature (PubMed), and microarray chips (screening of differentially expressed genes using limma package in R, |log2 fold change (FC)| > 1). Further, protein-protein interactions (PPIs) were investigated through the Human Protein Reference Database. miRNA-target relationships were also identified using the online software TargetScan. Currently, HPCDb contains 3284 genes, 120 miRNAs, 589 SNPs, 10,139 PPIs, and 3904 miRNA-target pairs. The detailed information on PC-related genes (e.g., gene identifier (ID), symbol, synonyms, full name, chip sets, expression alteration, PubMed ID, and PPIs), miRNAs (e.g., accession number, chromosome location, related disease, PubMed ID, and miRNA-target interactions), and SNPs (e.g., SNP ID, allele, gene, PubMed ID, chromosome location, and disease) is presented through user-friendly query interfaces or convenient links to NCBI GEO, NCBI PubMed, NCBI Gene, NCBI dbSNP, and miRBase. Overall, HPCDb provides biologists with relevant information on human PC-related molecules at multiple levels, helping to generate new hypotheses or identify candidate markers.

bioinformatics

A Synthetic Microbial Operational Amplifier

Synthetic biology has created oscillators, latches, logic gates, logarithmically linear circuits, and load drivers that have electronic analogs in living cells. The ubiquitous operational amplifier, which allows circuits to operate robustly and precisely has not been built with bio-molecular parts. As in electronics, a biological operational-amplifier could greatly improve the predictability of circuits despite noise and variability, a problem that all cellular circuits face. Here, we show how to create a synthetic 3-stage inducer-input operational amplifier with a differential transcription-factor stage, a CRISPR-based push-pull stage, and an enzymatic output stage with just 5 proteins including dCas9. Our Bio-OpAmp expands the toolkit of fundamental circuits available to bioengineers or biologists, and may shed insight into biological systems that require robust and precise molecular homeostasis and regulation.\n\nOne Sentence SummaryA synthetic bio-molecular operational amplifier that can enable robust, precise, and programmable homeostasis and regulation in living cells with just 5 protein parts is described.

synthetic biology

GEM: A manifold learning based framework for reconstructing spatial organizations of chromosomes

Decoding the spatial organizations of chromosomes has crucial implications for studying eukaryotic gene regulation. Recently, Chromosomal conformation capture based technologies, such as Hi-C, have been widely used to uncover the interaction frequencies of genomic loci in high-throughput and genome-wide manner and provide new insights into the folding of three-dimensional (3D) genome structure. In this paper, we develop a novel manifold learning framework, called GEM (Genomic organization reconstructor based on conformational Energy and Manifold learning), to elucidate the underlying 3D spatial organizations of chromosomes from Hi-C data. Unlike previous chromatin structure reconstruction methods, which explicitly assume specific relationships between Hi-C interaction frequencies and spatial distances between distal genomic loci, GEM is able to reconstruct an ensemble of chromatin conformations by directly embedding the neigh-boring affinities from Hi-C space into 3D Euclidean space based on a manifold learning strategy that considers both the fitness of Hi-C data and the biophysical feasibility of the modeled structures, which are measured by the conformational energy derived from our current biophysical knowledge about the 3D polymer model. Extensive validation tests on both simulated interaction frequency data and experimental Hi-C data of yeast and human demonstrated that GEM not only greatly outperformed other state-of-art modeling methods but also reconstructed accurate chromatin structures that agreed well with the hold-out or independent Hi-C data and sparse geometric restraints derived from the previous fluorescence in situ hybridization (FISH) studies. In addition, as GEM can generate accurate spatial organizations of chromosomes by integrating both experimentally-derived spatial contacts and conformational energy, we for the first time extended our modeling method to recover long-range genomic interactions that are missing from the original Hi-C data. All these results indicated that GEM can provide a physically and physiologically valid 3D representations of the organizations of chromosomes and thus serve as an effective and useful genome structure reconstructor.

bioinformatics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics

Phosphate Is The Third Nutrient Monitored By TOR In Candida albicans And Provides A Target For Fungal-Specific Indirect TOR Inhibition

The TOR pathway regulates morphogenesis and responses to host cells in the fungal pathogen Candida albicans. Eukaryotic TOR complex 1 (TORC1) induces growth and proliferation in response to nitrogen and carbon source availability. Our unbiased genetic approach seeking new components of TORC1 signaling in C. albicans revealed that the phosphate transporter Pho84 is required for normal TORC1 activity. We found that mutants in PHO84 are hypersensitive to rapamycin and, in response to phosphate feeding, generate less phosphorylated ribosomal protein S6 (P-S6) than wild type. The small GTPase Gtr1, a component of the TORC1-activating EGO complex, links Pho84 to TORC1. Mutants in Gtr1, but not in another TORC1-activating GTPase, Rhb1, are defective in the P-S6 response to phosphate. Overexpression of Gtr1 and of a constitutively active Gtr1Q67L mutant suppress TORC1-related defects. In S. cerevisiae pho84 mutants, constitutively active Gtr1 suppresses a TORC1 signaling defect but does not rescue rapamycin hypersensitivity. Hence connections from phosphate homeostasis to TORC1 may differ between C. albicans and S. cerevisiae. The converse direction of signaling, from TORC1 to the phosphate homeostasis (PHO) regulon, previously observed in S. cerevisiae, was genetically demonstrated in C. albicans using conditional TOR1 alleles. A small molecule inhibitor of Pho84, an FDA-approved drug, inhibits TORC1 signaling and potentiates the activity of the antifungals amphotericin B and micafungin. Anabolic TORC1-dependent processes require significant amounts of phosphate. Our study demonstrates that phosphate availability is monitored and also controlled by TORC1, and that TORC1 can be indirectly targeted by inhibiting Pho84.\n\nSignificanceThe human fungal pathogen Candida albicans uses the TOR signaling pathway to contend with varying host environments and thereby regulate cell growth. Seeking novel components of the C. albicans TOR pathway we identified a cell-surface phosphate importer, Pho84, and its molecular link to TOR complex 1 (TORC1). Since phosphorus is a critical element for anabolic processes like DNA replication, ribosome biogenesis, translation and membrane biosynthesis, TORC1 monitors its availability in regulating these processes. By depleting the central kinase in the TORC1 pathway, we showed that TORC1 signaling modulates regulation of phosphate acquisition. An FDA-approved small-molecule inhibitor of Pho84 inhibits TORC1 signaling and potentiates the activity of the gold-standard antifungal amphotericin B and the echinocandin micafungin.

microbiology

Metagenomic binning through low density hashing

Bacterial microbiomes of incredible complexity are found throughout the world, from exotic marine locations to the soil in our yards to within our very guts. With recent advances in Next-Generation Sequencing (NGS) technologies, we have vastly greater quantities of microbial genome data, but the nature of environmental samples is such that DNA from different species are mixed together. Here, we present Opal for metagenomic binning, the task of identifying the origin species of DNA sequencing reads. Our Opal method introduces low-density, even-coverage hashing to bioinformatics applications, enabling quick and accurate metagenomic binning. Our tool is up to two orders of magnitude faster than leading alignment-based methods at similar or improved accuracy, allowing computational tractability on large metagenomic datasets. Moreover, on public benchmarks, Opal is substantially more accurate than both alignment-based and alignment-free methods (e.g. on SimHC20.500, Opal achieves 95% F1-score while Kraken and CLARK achieve just 91% and 88%, respectively); this improvement is likely due to the fact that the latter methods cannot handle computationally-costly long-range dependencies, which our even-coverage, low-density fingerprints resolve. Notably, capturing these long-range dependencies drastically improves Opals ability to detect unknown species that share a genus or phylum with known bacteria. Additionally, the family of hash functions Opal uses can be generalized to other sequence analysis tasks that rely on k-mer based methods to encode long-range dependencies.

bioinformatics

Characterizing RNA Pseudouridylation By Convolutional Neural Networks

The most prevalent post-transcriptional RNA modification, pseudouridine ({Psi}), also known as the fifth ribonucleoside, is widespread in rRNAs, tRNAs, snRNAs, snoRNAs and mRNAs. Pseudouridines in RNAs are implicated in many aspects of post-transcriptional regulation, such as the maintenance of translation fidelity, control of RNA stability and stabilization of RNA structure. However, our understanding of the functions, mechanisms as well as precise distribution of pseudourdines (especially in mRNAs) still remains largely unclear. Though thousands of RNA pseudouridylation sites have been identified by high-throughput experimental techniques recently, the landscape of pseudouridines across the whole transcriptome has not yet been fully delineated. In this study, we present a highly effective model, called PULSE (PseudoUridyLation Sites Estimator), to predict novel {Psi} sites from large-scale profiling data of pseudouridines and characterize the contextual sequence features of pseudouridylation. PULSE employs a deep learning framework, called convolutional neural network (CNN), which has been successfully and widely used for sequence pattern discovery in the literature. Our extensive validation tests demonstrated that PULSE can outperform conventional learning models and achieve high prediction accuracy, thus enabling us to further characterize the transcriptome-wide landscape of pseudouridine sites. Overall, PULSE can provide a useful tool to further investigate the functional roles of pseudouridylation in post-transcriptional regulation.

bioinformatics