Search bioRxivSearch

Biology subjects

Chen, K.

Publications and source records attributed to Chen, K..

At least 19 recordsLinked to original sources

The Dynamic Conformational Landscapes of the Protein Methyltransferase SETD8

Elucidating conformational heterogeneity of proteins is essential for understanding protein functions and developing exogenous ligands for chemical perturbation. While structural biology methods can provide atomic details of static protein structures, these approaches cannot in general resolve less populated, functionally relevant conformations and uncover conformational kinetics. Here we demonstrate a new paradigm for illuminating dynamic conformational landscapes of target proteins. SETD8 (Pr-SET7/SET8/KMT5A) is a biologically relevant protein lysine methyltransferase for in vivo monomethylation of histone H4 lysine 20 and nonhistone targets. Utilizing covalent chemical inhibitors and depleting native ligands to trap hidden high-energy conformational states, we obtained diverse novel X-ray structures of SETD8. These structures were used to seed massively distributed molecular simulations that generated six milliseconds of trajectory data of SETD8 in the presence or absence of its cofactor. We used an automated machine learning approach to reveal slow conformational motions and thus distinct conformational states of SETD8, and validated the resulting dynamic conformational landscapes with multiple biophysical methods. The resulting models provide unprecedented mechanistic insight into how protein dynamics plays a role in SAM binding and thus catalysis, and how this function can be modulated by diverse cancer-associated mutants. These findings set up the foundation for revealing enzymatic mechanisms and developing inhibitors in the context of conformational landscapes of target proteins.

biophysics

An oxide transport chain essential for balanced insulin signaling

Patients with overnutrition, obesity, the atherometabolic syndrome, and type 2 diabetes mellitus exhibit imbalanced insulin action, also called pathway-selective insulin resistance. To control glycemia, they require hyperinsulinemia that then overdrives ERK and hepatic de-novo lipogenesis. We recently reported that NADPH oxidase-4 regulates balanced insulin action. Here, we show that NADPH oxidase-4 is part of a new limb of insulin signaling that we abbreviate \"NSAPP\" after its five major proteins. The NSAPP pathway is an oxide transport chain that begins when insulin stimulates NADPH oxidase-4 to generate [Formula]. NADPH oxidase-4 hands [Formula] to superoxide dismutase-3 for conversion into H2O2. The pathway ends when aquaporin-3 channels H2O2 across the membrane to inactivate PTEN. Disruption of any component of the NSAPP chain, from NADPH oxidase-4 up to PTEN, leaves PTEN persistently active, thereby producing the same deadly pattern of imbalanced insulin action seen clinically. Unraveling the molecular basis for NSAPP dysfunction in overnutrition has now become a top priority.

cell biology

Assessment of population differentiation and linkage disequilibrium in Solanum pimpinellifolium using genome-wide high-density SNP markers

To mine new favorable alleles for tomato breeding, we investigated the feasibility of utilizing Solanum pimpinellifolium as a diverse panel of genome-wide association study through the restriction site-associated DNA sequencing technique. Previous attempts to conduct genome-wide association study using S. pimpinellifolium were impeded by an inability to correct for population stratification and by lack of high-density markers to address the issue of rapid linkage disequilibrium decay. In the current study, a set of 24,330 SNPs was identified using 99 S. pimpinellifolium accessions from the Tomato Genetic Resource Center. Approximately 84% PstI site-associated DNA sequencing regions were located in the euchromatic regions, resulting in the tagging of most SNPs on or near genes. Our genotypic data suggested that the optimum number of S. pimpinellifolium ancestral subpopulations was three, and accessions were classified into seven groups. In contrast to the SolCAP SNP genotypic data of previous studies, our SNP genotypic data consistently confirmed the population differentiation, achieving a relatively uniform correction of population stratification. Moreover, as expected, rapid linkage disequilibrium decay was observed in S. pimpinellifolium, especially in euchromatic regions. Approximately two-thirds of the flanking SNP markers did not display linkage disequilibrium. Our result suggests that higher density of molecular markers and more accessions are required to conduct the genome-wide association study utilizing the Solanum pimpinellifolium collection.

genetics

SiCloneFit: Bayesian inference of population structure, genotype, and phylogeny of tumor clones from single-cell genome sequencing data

Accumulation and selection of somatic mutations in a Darwinian framework result in intra-tumor heterogeneity (ITH) that poses significant challenges to the diagnosis and clinical therapy of cancer. Identification of the tumor cell populations (clones) and reconstruction of their evolutionary relationship can elucidate this heterogeneity. Recently developed single-cell DNA sequencing (SCS) technologies promise to resolve ITH to a single-cell level. However, technical errors in SCS datasets, including false-positives (FP), false-negatives (FN) due to allelic dropout and cell doublets, significantly complicate these tasks. Here, we propose a non-parametric Bayesian method that reconstructs the clonal populations as clusters of single cells, genotypes of each clone and the evolutionary relationships between the clones. It employs a tree-structured Chinese restaurant process as the prior on the number and composition of clonal populations. The evolution of the clonal populations is modeled by a clonal phylogeny and a finite-site model of evolution to account for potential mutation recurrence and losses. We probabilistically account for FP and FN errors, and cell doublets are modeled by employing a Beta-binomial distribution. We develop a Gibbs sampling algorithm comprising of partial reversible-jump and partial Metropolis-Hastings updates to explore the joint posterior space of all parameters. The performance of our method on synthetic and experimental datasets suggests that joint reconstruction of tumor clones and clonal phylogeny under a finite-site model of evolution leads to more accurate inferences. Our method is the first to enable this joint reconstruction in a fully Bayesian framework, thus providing measures of support of the inferences it makes.

genomics

BAMM-SC: A Bayesian mixture model for clustering droplet-based single cell transcriptomic data from population studies

The recently developed droplet-based single cell transcriptome sequencing (scRNA-seq) technology makes it feasible to perform a population-scale scRNA-seq study, in which the transcriptome is measured for tens of thousands of single cells from multiple individuals. Despite the advances of many clustering methods, there are few tailored methods for population-scale scRNA-seq studies. Here, we have developed a BAyesiany Mixture Model for Single Cell sequencing (BAMM-SC) method to cluster scRNA-seq data from multiple individuals simultaneously. Specifically, BAMM-SC takes raw data as input and can account for data heterogeneity and batch effect among multiple individuals in a unified Bayesian hierarchical model framework. Results from extensive simulations and application of BAMM-SC to in-house scRNA-seq datasets using blood, lung and skin cells from humans or mice demonstrated that BAMM-SC outperformed existing clustering methods with improved clustering accuracy and reduced impact from batch effects. BAMM-SC has been implemented in a user-friendly R package with a detailed tutorial available on www.pitt.edu/~Cwec47/singlecell.html.

bioinformatics

Dynamic plant height QTL revealed in maize through remote sensing phenotyping using a high-throughput unmanned aerial vehicle (UAV)

Plant height is the key factor for plant architecture, biomass and yield in maize (Zea mays). In this study, plant height was investigated using unmanned aerial vehicle high-throughput phenotypic platforms (UAV-HTPPs) for maize diversity inbred lines at four important growth stages. Using an automated pipeline, we extracted accurate plant heights. We found that in temperate regions, from sowing to the jointing period, the growth rate for temperate maize was faster than tropical maize. However, from jointing to flowering stage, tropical maize maintained a vigorous growth state, and finally resulted in a taller plant than temperate lines. Genome-wide association study for temperate, tropical and both groups identified a total of 238 quantitative trait locus (QTLs) for the 16 plant height related traits over four growth periods. And, we found that plant height at different stages were controlled by different genes, for example, PIN1 controlled plant height at the early stage and PIN11 at the flowering stages. In this study, the plant height data collected by the UAV-HTTPs were credible and the genetic mapping power is high, indicating that the application of this UAV-HTTPs into the study of plant height will have great prospects.\n\nHighlightWe used UAV-based sensing platform to investigate plant height over 4 growth stages for different maize populations, and detected numbers of reliable QTLs using GWAS.

genetics

Sequencing of Panax notoginseng genome reveals genes involved in disease resistance and ginsenoside biosynthesis

Panax notoginseng is a traditional Chinese herb with high medicinal and economic value. There has been considerable research on the pharmacological activities of ginsenosides contained in Panax spp.; however, very little is known about the ginsenoside biosynthetic pathway. We reported the first de novo genome of 2.36 Gb of sequences from P. notoginseng with 35,451 protein-encoding genes. Compared to other plants, we found notable gene family contraction of disease-resistance genes in P. notoginseng, but notable expansion for several ATP-binding cassette (ABC) transporter subfamilies, such as the Gpdr subfamily, indicating that ABCs might be an additional mechanism for the plant to cope with biotic stress. Combining eight transcriptomes of roots and aerial parts, we identified several key genes, their transcription factor binding sites and all their family members involved in the synthesis pathway of ginsenosides in P. notoginseng, including dammarenediol synthase, CYP716 and UGT71. The complete genome analysis of P. notoginseng, the first in genus Panax, will serve as an important reference sequence for improving breeding and cultivation of this important nutraceutical and medicinal but vulnerable plant species.

genomics

Microbial coexistence through chemical-mediated interactions

Many microbial functions happen within communities of interacting species. Explaining how species with intrinsically disparate fitness can coexist is important for applications such as manipulating host-associated microbiota or engineering industrial communities. Previous coexistence studies have often neglected interaction mechanisms. Here, we formulate and experimentally constrain a model in which chemical mediators of microbial interactions (e.g. metabolites or waste-products) are explicitly incorporated. We construct many instances of coexistence by simulating community assembly through enrichment and ask how species interactions can explain coexistence. We show that growth-facilitating influences between members are favored in assembled communities. Among negative influences, self-restraint, such as production of self-inhibiting waste, contributes to coexistence, whereas inhibition of other species disrupts coexistence. Coexistence is also favored when interactions are mediated by depletable chemicals that get consumed or degraded, rather than by reusable chemicals that are unaffected by recipients. Our model creates null predictions for coexistence driven by chemical-mediated interactions.

ecology

SCMarker: ab initio marker selection for single cell transcriptome profiling

Single-cell RNA-sequencing data generated by a variety of technologies, such as Drop-seq and SMART-seq, can reveal simultaneously the mRNA transcript levels of thousands of genes in thousands of cells. It is often important to identify informative genes or cell-type-discriminative markers to reduce dimensionality and achieve informative cell typing results. We present an ab initio method that performs unsupervised marker selection by identifying genes that have subpopulation-discriminative expression levels and are co- or mutually-exclusively expressed with other genes. Consistent improvements in cell-type classification and biologically meaningful marker selection are achieved by applying SCMarker on various datasets in multiple tissue types, followed by a variety of clustering algorithms. The source code of SCMarker is publicly available at https://github.com/KChen-lab/SCMarker.\n\nAuthor SummarySingle cell RNA-sequencing technology simultaneously provides the mRNA transcript levels of thousands of genes in thousands of cells. A frequent requirement of single cell expression analysis is the identification of markers which may explain complex cellular states or tissue composition. We propose a new marker selection strategy (SCMarker) to accurately delineate cell types in single cell RNA-sequencing data by identifying genes that have bi/multi-modally distributed expression levels and are co- or mutually-exclusively expressed with some other genes. Our method can determine the cell-type-discriminative markers without referencing to any known transcriptomic profiles or cell ontologies, and consistently achieves accurate cell-type-discriminative marker identification in a variety of scRNA-seq datasets.

bioinformatics

Systematic discovery of uncharacterized transcription factors in Escherichia coli K-12 MG1655

Transcriptional regulation enables cells to respond to environmental changes. Yet, among the estimated 304 candidate transcription factors (TFs) in Escherichia coli K-12 MG1655, 185 have been experimentally identified and only a few tens of them have been fully characterized by ChIP methods. Understanding the remaining TFs is key to improving our knowledge of the E. coli transcriptional regulatory network (TRN). Here, we developed an integrated workflow for the computational prediction and comprehensive experimental validation of TFs using a suite of genome-wide experiments. We applied this workflow to: 1) identify 16 candidate TFs from over a hundred candidate uncharacterized genes; 2) capture a total of 255 DNA binding peaks for 10 candidate TFs resulting in six high-confidence binding motifs; 3) reconstruct the regulons of these 10 TFs by determining gene expression changes upon deletion of each TF; and 4) determine the regulatory roles of three TFs (YiaJ, YdcI, and YeiE) as regulators of L-ascorbate utilization, proton transfer and acetate metabolism, and iron homeostasis under iron limited condition, respectively. Together, these results demonstrate how this workflow can be used to discover, characterize, and elucidate regulatory functions of uncharacterized TFs in parallel.

microbiology

Disruption of cortical dopaminergic modulation delays licking initiation

Dysfunction of motor cortices is thought to contribute to motor disorders such as Parkinsons disease (PD). However, little is known on the link between cortical dopaminergic loss, abnormalities in motor cortex neural activity and motor deficits. We address the role of dopamine in modulating motor cortical activity by focusing on the anterior lateral motor cortex (ALM) of mice performing a cued-licking task. We first demonstrate licking deficits and concurrent alterations of spiking activity in ALM of mice with unilateral depletion of dopaminergic neurons (i.e., mice injected with 6-OHDA into the medial forebrain bundle). Hemi-lesioned mice displayed delayed licking initiation, shorter duration of licking bouts, and lateral deviation of tongue protrusions. In parallel with these motor deficits, we observed a reduction in the prevalence of cue responsive neurons and altered preparatory activity. Acute and local blockade of D1 receptors in ALM recapitulated some of the key behavioral and neural deficits observed in hemi-lesioned mice. Altogether, our data show a direct relationship between cortical D1 receptor modulation, cue-evoked and preparatory activity in ALM, and licking initiation.\n\nSIGNIFICANCE STATEMENTThe link between dopaminergic signaling, motor cortical activity and motor deficits is not fully understood. This manuscript describes alterations in neural activity of the anterior lateral motor cortex (ALM) that correlate with licking deficits in mice with unilateral dopamine depletion or with intra-ALM infusion of dopamine antagonist. The findings emphasize the importance of cortical dopaminergic modulation in motor initiation. These results will appeal not only to researchers interested in cortical control of licking, but also to a broader audience interested in motor control and dopaminergic modulation in physiological and pathological conditions. Specifically, our data suggest that dopamine deficiency in motor cortex could play a role in the pathogenesis of the motor symptoms of Parkinsons disease.

neuroscience

Dynamics of the sex ratio in Tetrahymena thermophila

Sex is often hailed as one of the major successes in evolution, and in sexual organisms the maintenance of proper sex ratio is crucial. As a large unicellular eukaryotic lineage, ciliates exhibit tremendous variation in mating systems, especially the number of sexes and the mechanism of sex determination (SD), and yet how the populations maintain proper sex ratio is poorly understood. Here Tetrahymena thermophila, a ciliate with seven mating types (sexes) and probabilistic SD mechanism, is analyzed from the standpoint of population genetics. It is found based on a newly developed population genetics model that there are plenty of opportunities for both the co-existence of all seven sexes and the fixation of a single sex, pending on several factors, including the strength of natural selection. To test the validity of predictions, five experimental populations of T. thermophila were maintained in the laboratory so that the factors that can influence the dynamics of sex ratio could be controlled and measured. Furthermore, whole-genome sequencing was employed to examine the impact of newly arisen mutations. Overall, it is found that the experimental observations highly support theoretical predictions. It is expected that the newly established theoretical framework is applicable in principle to other multi-sex organisms to bring more insight into the understanding of the maintenance of multiple sexes in a natural population.

evolutionary biology

Exendin-4 disrupts responding to reward predictive incentive cues in rats

Exendin-4 (EX4) is a GLP-1 receptor agonist used clinically to control glycemia in Type-2 Diabetes Mellitus (T2DM), with the additional effect of promoting weight loss. The weight loss seen with EX4 is attributable to the varied peripheral and central effects of GLP-1, with contributions from the mesolimbic dopamine pathway that are implicated in cue-induced reward seeking. GLP-1 receptor agonists reduce preference for palatable foods (i.e. sweet and fat) as well as the motivation to obtain and consume these foods. Accumulating evidence suggest that GLP-1 receptor activity can attenuate cue-induced reward seeking behaviors. In the present study, we tested the effects of EX4 (0.6, 1.2, and 2.4 {micro}g/kg i.p.) on incentive cue (IC) responding. This rat model required rats to emit a nosepoke response during an intermittent audiovisual cue to obtain a sucrose reward (10% solution). EX4 dose-dependently attenuated responding to reward predictive cues, and increased latencies of the cue response and reward cup entry to consume the sucrose reward. Moreover, EX4 dose-dependently decreased the number of nosepokes relative to the number of cue presentations during the session. There was no drug effect on the number of reward cup entries per reward earned during the session, a related reward-seeking behavior with similar locomotor demand. Interestingly, there was a dose-dependent effect of time on the responding to reward predictive cues and nosepoke response latency, such that 2.4 {micro}g/kg of EX4 delayed responding to the initial IC of the behavioral session. Together, these findings suggest that agonism of the GLP-1 receptor with EX4 modulates the incentive properties of cues attributed with motivational significance.

pharmacology and toxicology

Ankyrin-G Regulates Forebrain Connectivity and Network Synchronization via Interaction with GABARAP

GABAergic circuits are critical for the synchronization and higher order function of brain networks, and defects in this circuitry are linked to neuropsychiatric diseases, including bipolar disorder, schizophrenia, and autism. Work in cultured neurons has shown that ankyrin-G plays a key role in the regulation of GABAergic synapses on the axon initial segment and somatodendritic domain of pyramidal neurons where it interacts directly with the GABAO_SCPCAPAC_SCPCAP receptor associated protein (GABARAP) to stabilize cell surface GABAO_SCPCAPAC_SCPCAP receptors. Here, we generated a knock-in mouse model expressing a mutation that abolishes the ankyrin-G/GABARAP interaction (Ank3 W1989R) to understand how ankyrin-G and GABARAP regulate GABAergic circuitry in vivo. We found that Ank3 W1989R mice exhibit a striking reduction in forebrain GABAergic synapses resulting in pyramidal cell hyperexcitability and disruptions in network synchronization. In addition, we identified changes in pyramidal cell dendritic spines and axon initial segments consistent with compensation for hyperexcitability. Finally, we identified the ANK3 W1989R variant in a family with bipolar disorder, suggesting a potential role of this variant in disease. Our results highlight the importance of ankyrin-G in regulating forebrain circuitry and provide novel insights into how ANK3 loss-of-function variants may contribute to human disease.

neuroscience

An inter-chromosomal transcription hub activates the unfolded protein response in plasma cells

Previous studies have indicated that the transcription signature of antibody-secreting cells is closely associated with the induction of the unfolded protein response pathway (UPR). Here we have used genome-wide and single cell analyses to examine the folding patterns of plasma cell genomes. We found that plasma cells adopt a cartwheel configuration and undergo large-scale changes in chromatin folding at genomic regions associated with a plasma cell specific transcription signature. During plasma cell differentiation, Blimp1 assembles into an inter-chromosomal transcription hub with genes associated with the UPR, biosynthesis of the endoplasmic reticulum (ER) as well as a cluster of genes linked with Alzheimers disease. We suggest that the assembly of the Blimp1-UPR-ER transcription hub permits the coordinate activation of a wide spectrum of genes that collectively establish plasma cell identity.

immunology

Recombinant expression of Proteorhodopsin and biofilm regulators in Escherichia coli for nanoparticle binding and removal in a wastewater treatment model

The small size of nanoparticles is both an advantage and a problem. Their high surface-area-to-volume ratio enables novel medical, industrial, and commercial applications. However, their small size also allows them to evade conventional filtration during water treatment, posing health risks to humans, plants, and aquatic life. This project aims to remove nanoparticles during wastewater treatment using genetically modified Escherichia coli in two ways: 1) binding citrate-capped nanoparticles with the membrane protein Proteorhodopsin, and 2) trapping nanoparticles using Escherichia coli biofilm produced by overexpressing two regulators: OmpR234 and CsgD. We demonstrate experimentally that Escherichia coli expressing Proteorhodopsin binds to 60 nm citrate-capped silver nanoparticles. We also successfully upregulate biofilm production and show that Escherichia coli biofilms are able to trap 30 nm gold particles. Finally, both Proteorhodopsin and biofilm approaches are able to bind and remove nanoparticles in simulated wastewater treatment tanks. We envision integrating our trapping system in both rural and urban wastewater treatment plants to efficiently capture all nanoparticles before treated water is released into the environment.\n\nFinancial DisclosureThis work was funded by the Taipei American School. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.\n\nCompeting InterestsThe authors have declared that no competing interests exist.\n\nEthics StatementN/A\n\nData AvailabilityYes - all data are fully available without restriction. Sequences for the plasmids used in this study are available through the Registry of Standard Biological Parts. Links to raw data are included in Supplementary Information.

synthetic biology

Multi-scale model of the proteomic and metabolic consequences of reactive oxygen species

Catalysis using iron-sulfur clusters and transition metals can be traced back to the last universal common ancestor. The damage to metalloproteins caused by reactive oxygen species (ROS) can completely inhibit cell growth when unmanaged and thus elicits an essential stress response that is universal and fundamental in biology. We develop a computable multi-scale description of the ROS stress response in Escherichia coli. We show that this quantitative framework allows for the understanding and prediction of ROS stress responses at three levels: 1) pathways: amino acid auxotrophies, 2) networks: the systemic response to ROS stress, and 3) genetic basis: adaptation to ROS stress during laboratory evolution. These results show that we can now develop fundamental and quantitative genotype-phenotype relationships for stress responses on a genome-wide basis.

systems biology

Combining accurate tumour genome simulation with crowd sourcing to benchmark somatic structural variant detection

BackgroundThe phenotypes of cancer cells are driven in part by somatic structural variants. Structural variants can initiate tumors, enhance their aggressiveness and provide unique therapeutic opportunities. Whole-genome sequencing of tumors can allow exhaustive identification of the specific structural variants present in an individual cancer, facilitating both clinical diagnostics and the discovery of novel mutagenic mechanisms. A plethora of somatic structural variant detection algorithms have been created to enable these discoveries, however there are no systematic benchmarks of them. Rigorous performance evaluation of somatic structural variant detection methods has been challenged by the lack of gold-standards, extensive resource requirements and difficulties arising from the need to share personal genomic information.\n\nResultsTo facilitate structural variant detection algorithm evaluations, we create a robust simulation framework for somatic structural variants by extending the BAMSurgeon algorithm. We then organize and enable a crowd-sourced benchmarking within the ICGC-TCGA DREAM Somatic Mutation Calling Challenge (SMC-DNA). We report here the results of structural variant benchmarking on three different tumors, comprising 204 submissions from 15 teams. In addition to ranking methods, we identify characteristic error-profiles of individual algorithms and general trends across them. Surprisingly, we find that ensembles of analysis pipelines do not always outperform the best individual method, indicating a need for new ways to aggregate somatic structural variant detection approaches.\n\nConclusionsThe synthetic tumors and somatic structural variant detection leaderboards remain available as a community benchmarking resource, and BAMSurgeon is available at https://github.com/adamewing/bamsurgeon.

bioinformatics