Search bioRxivSearch

Biology subjects

Song, J.

Publications and source records attributed to Song, J..

At least 19 recordsLinked to original sources

The filopodial scaffold polyphosphate dictates cell adhesion-versus-invasion decisions

Inorganic polyphosphate (polyP) is an ancient polymer conserved across all life, serving cell type and location specific functions in every major compartment. Yet its role at the plasma membrane, where it accumulates to peak levels in many primary cells, is largely unknown. Here we identify polyP as a stabilizing component of filopodia, actin based membrane protrusions that govern cell adhesion, contact inhibition, and chemotaxis. Elevating cellular polyP increases filopodial stability and enhances cell adhesion, whereas reducing polyP accelerates filopodial disassembly and promotes cell migration. Mechanistically, we find that polyP acts as a structural filopodial scaffold, recruiting and organizing IRSp53, a membrane curvature inducing protein. We show that metastatic fibroblasts and breast cancer organoids carry markedly reduced and intracellularly reorganized polyP levels relative to their non transformed counterparts. Restoring endogenous polyP via lipid nanoparticle delivery suppresses their invasive phenotypes and reverses prometastatic gene expression signatures, implicating polyP as a primordial tumor suppressor.

cell biology

Multilayered mechanisms ensure that short chromosomes recombine in meiosis

To segregate accurately during meiosis, homologous chromosomes in most species must recombine. Very small chromosomes would risk missegregation if recombination were randomly distributed, so the double-strand breaks (DSBs) that initiate recombination are not haphazard. How this nonrandomness is controlled is not under-stood. Here we demonstrate that Saccharomyces cerevisiae integrates multiple, temporally distinct pathways to regulate chromosomal binding of pro-DSB factors Rec114 and Mer2, thereby controlling duration of a DSB-competent state. Homologous chromosome engagement regulates Rec114/Mer2 dissociation late in prophase, whereas replication timing and proximity to centromeres or telomeres influence timing and amount of Rec114/Mer2 accumulation early. A distinct early mechanism boosts Rec114/Mer2 binding quickly to high levels specifically on the shortest chromosomes, dependent on chromosome axis proteins and subject to selection pressure to maintain hyperrecombinogenic properties of these chromosomes. Thus, an organisms karyotype and its attendant risk of meiotic missegregation influence the shape and evolution of its recombination landscape.

genetics

Molecular mimicry in deoxy-nucleotide catalysis: the structure of Escherichia coli dGTPase reveals the molecular basis of dGTP selectivity

Deoxynucleotide triphosphate triphosphyohydrolyases (dNTPases) play a critical role in cellular survival and DNA replication through the proper maintenance of cellular dNTP pools by hydrolyzing dNTPs into deoxynucleosides and inorganic triphosphate (PPPi). While the vast majority of these enzymes display broad activity towards canonical dNTPs, exemplified by Sterile Alpha Motif (SAM) and Histidine-aspartate (HD) domain-containing protein 1 (SAMHD1), which blocks reverse transcription of retroviruses in macrophages by maintaining dNTP pools at low levels, Escherichia coli (Ec)-dGTPase is the only known enzyme that specifically hydrolyzes dGTP. However, the mechanism behind dGTP selectivity is unclear. Here we present the free-, ligand (dGTP)- and inhibitor (GTP)-bound structures of hexameric E. coli dGTPase. To obtain these structures, we applied UV-fluorescence microscopy, video analysis and highly automated goniometer-based instrumentation to map and rapidly position individual crystals randomly-located on fixed target holders, resulting in the highest indexing-rates observed for a serial femtosecond crystallography (SFX) experiment. The structure features a highly dynamic active site where conformational changes are coupled to substrate (dGTP), but not inhibitor binding, since GTP locks dGTPase in its apo form. Moreover, despite no sequence homology, dGTPase and SAMHD1 share similar active site and HD motif architectures; however, dGTPase residues at the end of the substrate-binding pocket mimic Watson Crick interactions providing Guanine base specificity, while a 7 [A] cleft separates SAMHD1 residues from dNTP bases, abolishing nucleotide-type discrimination. Furthermore, the structures sheds light into the mechanism by which long distance binding (25 [A]) of single stranded DNA in an allosteric site primes the active site by conformationally \"opening\" a tyrosine gate allowing enhanced substrate binding.\n\nSignificance StatementdNTPases play a critical role in cellular survival through maintenance of cellular dNTP. While dNTPases display activity towards dNTPs, such as SAMHD1 -which blocks reverse transcription of HIV-1 in macrophages- Escherichia coli (Ec)-dGTPase is the only known enzyme that specifically hydrolyzes dGTP. Here we use novel free electron laser data collection to shed light into the mechanisms of (Ec)-dGTPase selectivity. The structure features a dynamic active site where conformational changes are coupled to dGTP binding. Moreover, despite no sequence homology between (Ec)-dGTPase and SAMHD1, both enzymes share similar active site architectures; however, dGTPase residues at the end of the substrate-binding pocket provide dGTP specificity, while a 7 [A] cleft separates SAMHD1 residues from dNTP.

biophysics

Genome-wide CNV study and functional evaluation identified CTDSPL as tumour suppressor gene for cervical cancer

We have investigated copy number variations (CNVs) in relation to cervical cancer by analyzing 731,422 single-nucleotide polymorphisms (SNPs) in 1,034 cervical cancer cases and 3,948 controls, followed by replication in 1,396 cases and 1,057 controls. We found that a 6367bp deletion in intron 1 of the CTD small phosphatase like gene (CTDSPL) was associated with 2.54-fold increased risk of cervical cancer (odds ratio =2.54, 95% confidence interval =2.08-3.12, P=2.0x10-19). This CNV is one of the strongest genetic risk variants identified so far for cervical cancer. The deletion removes the binding sites of zinc finger protein 263, binding protein 2 and interferon regulatory factor 1, and hence downregulates the transcription of CTDSPL. HeLa cells expressing CTDSPL showed a significant decrease in colony-forming ability. Compared with control groups, mice injected with HeLa cells expressing CTDSPL exhibited a significant reduction in tumour volume. Furthermore, CTDSPL-depleted immortalized End1/E6E7 could form tumours in NOD-SCID mice.

cancer biology

TMEM106B, a risk factor for FTLD and aging, has an intrinsically disordered cytoplasmic domain

TMEM106B was initially identified as a risk factor for FTLD, but recent studies highlighted its general role in neurodegenerative diseases. Very recently TMEM106B has also been characterized to regulate aging phenotypes. TMEM106B is a 274-residue lysosomal protein whose cytoplasmic domain functions in the endosomal/autophagy pathway by dynamically and transiently interacting with diverse categories of proteins but the underlying structural basis remains completely unknown. Here we conducted bioinformatics analysis and biophysical characterization by CD and NMR spectroscopy, and obtained results reveal that the TMEM106B cytoplasmic domain is intrinsically disordered with no well-defined three-dimensional structure. Nevertheless, detailed analysis of various multi-dimensional NMR spectra allowed defining residue-specific conformations and dynamics. Overall, the TMEM106B cytoplasmic domain is lacking of any tight tertiary packing and relatively flexible. However, several segments are populated with dynamic/nascent secondary structures and have relatively restricted backbone motions. In particular, the fragment Ser12-Met36 is highly populated with - helix conformation. Our study thus decodes that being intrinsically disordered allows the TMEM106B cytoplasmic domain to dynamically and transiently interact with a variety of distinct partners.

biophysics

Misfolded proteins share a common capacity in disrupting LLPS organizing membrane-less organelles

Profilin-1 mutants cause ALS by gain of toxicity but the underlying mechanism remains unknown. Here we showed that three PFN1 mutants have differential capacity in disrupting dynamics of FUS liquid droplets underlying the formation of stress granules (SGs). Subsequently we extensively characterized conformations, dynamics and hydrodynamic properties of C71G-PFN1, FUS droplets and their interaction by NMR spectroscopy. C71G-PFN1 co-exists between the folded (55.2%) and unfolded (44.8%) states undergoing exchanges at 11.7 Hz, while its unfolded state non-specifically interacts with FUS droplets. Results together lead to a model for dynamic droplets to recruit misfolded proteins, which functions seemingly at great cost: simple accumulation of misfolded proteins within liquid droplets is sufficient to reduce their dynamics. Further aggregation of misfolded proteins within droplets might irreversibly disrupt/destroy structures and dynamics of droplets, as increasingly observed on SGs, an emerging target for various neurodegenerative diseases. Therefore, our study implies that other misfolded proteins might also share the capacity in disrupting LLPS.

biophysics

Characterization of a human-specific tandem repeat associated with bipolar disorder and schizophrenia

Bipolar disorder (BD) and schizophrenia (SCZ) are highly heritable diseases that affect over 3% of individuals worldwide. Genomewide association studies have strongly and repeatedly linked risk for both of these neuropsychiatric diseases to a 100 kb interval in the third intron of the human calcium channel gene CACNA1C. However, the causative mutation is not yet known. We have identified a novel human-specific tandem repeat in this region that is composed of 30 bp units, often repeated hundreds of times. This large tandem repeat is unstable using standard polymerase chain reaction and bacterial cloning techniques, which may have resulted in its incorrect size in the human reference genome. The large 30-mer repeat region is polymorphic in both size and sequence in human populations. Particular sequence variants of the 30-mer are associated with risk status at several flanking single nucleotide polymorphisms in the third intron of CACNA1C that have previously been linked to BD and SCZ. The tandem repeat arrays function as enhancers that increase reporter gene expression in a human neural progenitor cell line. Different human arrays vary in the magnitude of enhancer activity, and the 30-mer arrays associated with increased psychiatric disease risk status have decreased enhancer activity. Changes in the structure and sequence of these arrays likely contribute to changes in CACNA1C function during human evolution, and may modulate neuropsychiatric disease risk in modern human populations.

genetics

Nucleolar stress triggers the irreversible cell cycle slow down leading to cell death during replicative aging in Saccharomyces cerevisiae

The accumulation of Extrachromosomal rDNA Circles (ERCs) and their asymmetric segregation upon division have been hypothesized to be responsible for replicative senescence in mother yeasts and rejuvenation in daughter cells. However, it remains unclear by which molecular mechanisms ERCs would trigger the irreversible cell cycle slow-down leading to cell death. We show that ERCs accumulation is concomitant with a nucleolar stress, characterized by a massive accumulation of pre-rRNAs in the nucleolus, leading to a loss of nucleus-to-cytoplasm ratio, decreased growth rate and cell-cycle slow-down. This nucleolar stress, observed in old mothers, is not inherited by rejuvenated daughters. Unlike WT, in the long-lived mutant fob1{triangleup}, a majority of cells is devoid of nucleolar stress and does not experience replicative senescence before death. Our study provides a unique framework to order the successive steps that govern the transition to replicative senescence and highlights the causal role of nucleolar stress in cellular aging.

cell biology

Two Active X-chromosomes Modulate the Growth, Pluripotency Exit and DNA Methylation Landscape of Mouse Naive Pluripotent Stem Cells through Different Pathways

During early mammalian development, the two X-chromosomes in female cells are active. Dosage compensation between XX female and XY male cells is then achieved by X-chromosome inactivation in female cells. Reprogramming female mouse somatic cells into induced pluripotent stem cells (iPSCs) leads to X-chromosome reactivation. The extent to which increased X-chromosome dosage (X-dosage) in female iPSCs leads to differences in the molecular and cellular properties of XX and XY iPSCs is still unclear. We show that chromatin accessibility in mouse iPSCs is modulated by X-dosage. Specific sets of transcriptional regulator motifs are enriched in chromatin with increased accessibility in XX or XY iPSCs. We show that the transcriptome, growth and pluripotency exit are also modulated by X-dosage in iPSCs. To understand the mechanisms by which increased X-dosage modulates the molecular and cellular properties of mouse pluripotent stem cells, we used heterozygous deletions of the X-linked gene Dusp9 in XX embryonic stem cells. We show that X-dosage regulates the transcriptome, open chromatin landscape, growth and pluripotency exit largely independently of global DNA methylation. Our results uncover new insights into X-dosage in pluripotent stem cells, providing principles of how gene dosage modulates the epigenetic and genetic mechanisms regulating cell identity.

developmental biology

Structural Capacitance in Protein Evolution and Human Diseases

Canonical mechanisms of protein evolution include the duplication and diversification of pre-existing folds through genetic alterations that include point mutations, insertions, deletions, and copy number amplifications, as well as post-translational modifications that modify processes such as folding efficiency and cellular localization. Following a survey of the human mutation database, we have identified an additional mechanism, that we term structural capacitance, which results in the de novo generation of microstructure in previously disordered regions. We suggest that the potential for structural capacitance confers select proteins with the capacity to evolve over rapid timescales, facilitating saltatory evolution as opoposed to exclusively canonical Darwinian mechanisms. Our results implicate the elements of protein microstructure generated by this distinct mechanism in the pathogenesis of a wide variety of human diseases. The benefits of rapidly furnishing the potential for evolutionary change conferred by structural capacitance are consequently counterbalanced by this accompanying risk, with the extent of this determined by the host immune system. The phenomenon of structural capacitance has implications ranging from the ancestral diversification of protein folds to the engineering of synthetic proteins with enhanced evolvability.

biochemistry

DeepGS: Predicting phenotypes from genotypes using Deep Learning

MotivationGenomic selection (GS) is a new breeding strategy by which the phenotypes of quantitative traits are usually predicted based on genome-wide markers of genotypes using conventional statistical models. However, the GS prediction models typically make strong assumptions and perform linear regression analysis, limiting their accuracies since they do not capture the complex, non-linear relationships within genotypes, and between genotypes and phenotypes.\n\nResultsWe present a deep learning method, named DeepGS, to predict phenotypes from genotypes. Using a deep convolutional neural network, DeepGS uses hidden variables that jointly represent features in genotypic markers when making predictions; it also employs convolution, sampling and dropout strategies to reduce the complexity of high-dimensional marker data. We used a large GS dataset to train DeepGS and compare its performance with other methods. In terms of mean normalized discounted cumulative gain value, DeepGS achieves an increase of 27.70%~246.34% over a conventional neural network in selecting top-ranked 1% individuals with high phenotypic values for the eight tested traits. Additionally, compared with the widely used method RR-BLUP, DeepGS still yields a relative improvement ranging from 1.44% to 65.24%. Through extensive simulation experiments, we also demonstrated the effectiveness and robustness of DeepGS for the absent of outlier individuals and subsets of genotypic markers. Finally, we illustrated the complementarity of DeepGS and RR-BLUP with an ensemble learning approach for further improving prediction performance.\n\nAvailabilityDeepGS is provided as an open source R package available at https://github.com/cma2015/DeepGS.

bioinformatics

PEA: an integrated R toolkit for plant epitranscriptome analysis

MotivationThe epitranscriptome, also known as chemical modifications of RNA (CMRs), is a newly discovered layer of gene regulation, the biological importance of which emerged through analysis of only a small fraction of CMRs detected by high-throughput sequencing technologies. Understanding of the epitranscriptome is hampered by the absence of computational tools for the systematic analysis of epitranscriptome sequencing data. In addition, no tools have yet been designed for accurate prediction of CMRs in plants, or to extend epitranscriptome analysis from a fraction of the transcriptome to its entirety.\n\nResultsHere, we introduce PEA, an integrated R toolkit to facilitate the analysis of plant epitranscriptome data. The PEA toolkit contains a comprehensive collection of functions required for read mapping, CMR calling, motif scanning and discovery, and gene functional enrichment analysis. PEA also takes advantage of machine learning technologies for transcriptome-scale CMR prediction, with high prediction accuracy, using the Positive Samples Only Learning algorithm, which addresses the two-class classification problem by using only positive samples (CMRs), in the absence of negative samples (non-CMRs). Hence PEA is a versatile epitranscriptome analysis pipeline covering CMR calling, prediction, and annotation, and we describe its application to predict N6-methyladenosine (m6A) modifications in Arabidopsis thaliana. Experimental results demonstrate that the toolkit achieved 71.6% sensitivity and 73.7% specificity, which is superior to existing m6A predictors. PEA is potentially broadly applicable to the in-depth study of epitranscriptomics.\n\nAvailabilityPEA is implemented using R and available at https://github.com/cma2015/PEA.

bioinformatics

The PTPRT pseudo-phosphatase domain is a denitrase

Protein tyrosine nitration occurs under both physiological and pathological conditions1. However, enzymes that remove this protein modification have not yet been identified. Here we report that the pseudo-phosphatase domain of protein tyrosine receptor T (PTPRT) is a denitrase that removes nitro-groups from tyrosine residues in paxillin. PTPRT normally functions as a tumor suppressor and is frequently mutated in a variety of human cancers including colorectal cancer2,3. We demonstrate that some of the tumor-derived mutations located in the pseudophosphatase domain impair the denitrase activity. Moreover, PTPRT mutant mice that inactivate the denitrase activity are susceptible to carcinogen-induced colon tumor formation. This study uncovers a novel enzymatic activity that is involved in tumor suppression.

cancer biology

Subtle perturbations of the maize methylome reveal genes and transposons silenced by DNA methylation

DNA methylation is a chromatin modification that can provide epigenetic regulation of gene and transposon expression. Plants utilize several pathways to establish and maintain DNA methylation in specific sequence contexts. The chromomethylase (CMT) genes maintain CHG (where H = A, C or T) methylation. The RNA-directed DNA methylation (RdDM) pathway is important for CHH methylation. Transcriptome analysis was performed in a collection of Zea mays lines carrying mutant alleles for CMT or RdDM-associated genes. While the majority of the transcriptome was not affected, we identified sets of genes and transposon families sensitive to context-specific decreases in DNA methylation in mutant lines. Many of the genes that are up-regulated in CMT mutant lines have high levels of CHG methylation, while genes that are differentially expressed in RdDM mutants are enriched for having nearby mCHH islands, providing evidence that context-specific DNA methylation directly regulates expression of a small number of genes. The analysis of a diverse set of inbred lines revealed that many genes regulated by CMTs exhibit natural variation for DNA methylation and gene expression. Transposon families with differential expression in the mutant genotypes show few defining features, though several families up-regulated in RdDM mutants show enriched expression in endosperm, highlighting the importance for this pathway during reproduction. Taken together, our findings suggest that while the number of genes and transposon families whose expression is reproducibly affected by mild perturbations in context-specific methylation is small, there are distinct patterns for loci impacted by RdDM and CMT mutants.

plant biology

NMR studies reveal that protein dynamics critically mediate aggregation of the well-folded and very soluble E. coli S1 ribosomal protein

Unlike mammalian aging associated with many hallmarks, E. coli aging is only significantly characterized by protein aggregation, thus offering an excellent model for addressing the relationship between protein aggregation and aging. Here we characterized conformations, unfolding and dynamics of ribosomal protein S1 and its D3/D5 domains using NMR, CD and fluorescence spectroscopy. S1 is a 557-residue modular protein containing six S1 motifs. Paradoxically, while S1 is well-folded and very soluble in vitro, it was found in various lists of aggregated E. coli proteins. Our results decipher: 1) S1 has dynamic inter-domain interactions. Strikingly, S1 and its D3/D5 domains have significantly exposed hydrophobic patches characterized by irreversible unfolding. 2) Although D5 has significantly restricted backbone motion on ps-ns time scale, it has global s-ms conformational dynamics and particularly high \"global breathing\" motions. 3) D5 assumes the conserved {beta}-barrel fold but contains large hydrophobic patches at least dynamically accessible. Taken together, our study reveals that S1 could be prone to aggregation due to significant dynamics at two levels: inter-domain interactions and individual domains, which may even render buried hydrophobic patches/cores accessible for driving aggregation. This mechanism is most likely to operate in many proteins of E. coli and other organisms including human.

biophysics

TDCPP exposure affects the concentrations of thyroid hormones in zebrafish

Previous studies show that TDCPP may interrupt the thyroid endocrine system, however, the potential mechanisms involved in these processes were largely unknown. In this study, zebrafish embryos/larvae were exposed to TDCPP until 120 hpf, by which time most of the organs of the larvae have completed development. In this study, the effects of TDCPP on HPT axis were examined and the thyroid hormone levels were measured after TDCPP treatment. Zebrafish (Danio rerio) embryos were treated with a series concentration of TDCPP (10, 20, 40, 80, 160 and 320 g/L) from 1 day post-fertilization (dpf) to 5 dpf. Exposure concentrations of TDCPP were determined based on the survival rates in each group. Total mRNA were isolated, first-strand cDNA were synthesis and qPCR were performed to detect the mRNA expression levels in hypothalamic-pituitary-thyroid (HPT) axis. The mRNA expression levels of genes involved in thyroid hormone homeostasis were increased in the TDCPP-treated larvae. The mRNA levels of genes involved in thyroid hormone synthesis were also increased in the embryos treated with TDCPP. Furthermore, exposure to TDCPP led to a dose-dependent effect on zebrafish development, including diminished hatching and survival rates, increased malformation. TDCPP treatment significantly reduced the T4 concentration in the 5 dpf zebrafish larvae, but increased the concentration of T3, suggesting the function of thyroid endocrine were interrupted in the TDCPP-exposed zebrafish. Taken together, these data indicated that TDCPP affected the thyroid hormone levels in the zebrafish larvae and could increased the mRNA expression levels of genes related to HPT axis, which further impaired the endocrine homeostasis and thyroid system.

pharmacology and toxicology

RRM domain of ALS/FTD-causing FUS interacts with membrane: an anchor of membraneless organelles to membranes?

526-residue FUS functions to self-assemble into reversible droplets/hydrogels, which could be further solidified into pathological fibrils. FUS is composed of N-terminal low-sequence complexity (LC); RNA-recognition motif (RRM) and C-terminal LC domains. FUS belongs to an emerging category of proteins which are capable of forming membraneless organelles in cells via phase separation. On the other hand, eukaryotic cells contain a large network of internal membrane systems. Therefore, it is of fundamental importance to address whether membraneless organelles can interact with membranes. Here we attempted to explore this by NMR HSQC titrations of three FUS domains with gradual addition of DMPC/DHPC bicelle, which mimics the bilayer membrane. We found that both N- and C-terminal LC domains showed no significant interaction with bicelle, but its well-folded RRM domain does dynamically interact with bicelle with an interface opposite to that for binding nucleic acids including RNA and ssDNA. If this in vitro observation also occurs in cells, to interact with membrane might represent a mechanism for dynamically organizing membraneless organelles to membranes to facilitate their physiological functions.

biophysics

TeachEnG: a Teaching Engine for Genomics

MotivationBioinformatics is a rapidly growing field that has emerged from the synergy of computer science, statistics, and biology. Given the interdisciplinary nature of bioinformatics, many students from diverse fields struggle with grasping bioinformatic concepts only from classroom lectures. Interactive tools for helping students reinforce their learning would be thus desirable. Here, we present an interactive online educational tool called TeachEnG (acronym for Teaching Engine for Genomics) for reinforcing key concepts in sequence alignment and phylogenetic tree reconstruction. Our instructional games allow students to align sequences by hand, fill out the dynamic programming matrix in the Needleman-Wunsch global sequence alignment algorithm, and reconstruct phylogenetic trees via the maximum parsimony and Unweighted Pair Group Method with Arithmetic mean (UPGMA) algorithms. With an easily accessible interface and instant visual feedback, TeachEnG will help promote active learning in bioinformatics.\n\nAvailability and Implementation: TeachEnG is freely available at http://song.igb.illinois.edu/TeachEnG/. It is written in JavaScript and compatible with Firefox, Safari, Chrome, and Microsoft Edge.\n\nContact: songi@illinois.edu

bioinformatics