Search bioRxivSearch

Biology subjects

Zhu, H.

Publications and source records attributed to Zhu, H..

At least 19 recordsLinked to original sources

Specificity of RNA folding and its association with evolutionarily adaptive mRNA secondary structure

Secondary structure is a fundamental feature for both noncoding and messenger RNA. However, our understandings about the secondary structure of mRNA, especially for the coding regions, remain elusive, likely due to translation and the lack of RNA binding proteins that sustain the consensus structure, such as those bind to noncoding RNA. Indeed, mRNA has recently been found to bear pervasive alternative structures, whose overall evolutionary and functional significance remained untested. We hereby approached this problem by estimating folding specificity, the probability that a fragment of RNA folds back to the same partner once re-folded. We showed that folding specificity for mRNA is lower than noncoding RNA, and displays moderate evolutionary conservation between orthologs and between paralogs. More importantly, we found that specific rather than alternative folding is more likely evolutionarily adaptive, since it is more frequently associated with functionally important genes or sites within a gene. Additional analysis in combination with ribosome density suggests the capability of modulating ribosome movement as one potential functional advantage provided by specific folding. Our findings revealed a novel facet of RNA structome with important functional and evolutionary implications, and points to a potential way of disentangling mRNA secondary structures maintained by natural selection from molecular noise.

evolutionary biology

Mitotic Regulators and the SHP2-MAPK Pathway Promote Insulin Receptor Endocytosis and Feedback Regulation of Insulin Signaling

Insulin controls glucose homeostasis and cell growth through bifurcated signaling pathways. Dysregulation of insulin signaling is linked to diabetes and cancer. The spindle checkpoint controls the fidelity of chromosome segregation during mitosis. Here, we show that insulin receptor substrate 1 and 2 (IRS1/2) cooperate with spindle checkpoint proteins to promote insulin receptor (IR) endocytosis through recruiting the clathrin adaptor complex AP2 to IR. A phosphorylation switch of IRS1/2 orchestrated by extracellularly regulated kinase 1 and 2 (ERK1/2) and Src homology phosphatase 2 (SHP2) ensures selective internalization of activated IR. SHP2 inhibition blocks this feedback regulation and growth-promoting IR signaling, prolongs insulin action on metabolism, and improves insulin sensitivity in mice. We propose that mitotic regulators and SHP2 promote feedback inhibition of IR, thereby limiting the duration of insulin signaling. Targeting this feedback inhibition can improve insulin sensitivity.

cell biology

Diagnostic Whole Exome Sequencing in Patients with Short Stature

Short stature is among the most common reasons for children being referred to the pediatric endocrinology clinics. The cause of short stature is broad, in which genetic factors play a substantial role, especially in primary growth disorders. However, identifying the molecular causes for short stature remains as a challenge because of the high heterogeneity of the phenotypes. Here, whole exome sequencing (WES) was used to identify the genetic causes of short stature with unknown etiology for 20 patients aged from 1 to 16 years old. The genetic causes of short stature were identified in 9 of the 20 patients, corresponding to a molecular diagnostic rate of 45%. Notably, in 2 of the 9 patients identified with genetic causes, the diagnosed diseases based on WES are different from the original clinical diagnosis. Our results highlight the clinical utility of WES in the diagnosis of rare, high heterogeneity disorders.

genetics

SCDT: Detecting CNVs of low chimeric ratio in cf-DNA

MotivationSequencing of cell-free DNA (cf-DNA) has enabled Noninvasive Prenatal Testing (NIPT) and\"liquid biopsy\" of cancers. However, while the aneuploidy and point mutations were focused on by most of NITP and liquid biopsy studies, detecting sub-chromosome CNVs that affect a few to dozens of megabases was rarely reported, likely attributable to the difficulty in accurately identifying them, especially for those present in a small fraction of cf-DNA.\n\nResultsWe developed a somatic CNV detection tool (SCDT), for detecting sub-chromosome CNVs in cf-DNA using whole genome sequencing (WGS) data or off-target reads in target sequencing data. Additional to using control samples for correcting genome position specific bias, two GC correction steps were performed, which regressed GC content of DNA fragments and that of genome bins, respectively. After GC correction, the coefficients of variation of copy ratios approximated the lower boundary of theoretical values, suggesting removing of almost all systematic errors. Finally, CNVs were detected by a piecewise least squares fitting based segmentation algorithm, which outperformed other segmentation methods. We applied SCDT on simulated and real maternal plasma samples, and target cf-DNA sequencing of 118 normal individuals and 240 cancer patients, and demonstrated high sensitivity and specificity.\n\nAvailabilitySCDT is available at https://github.com/Martiantian/Somatic_cnv_detect_tool.\n\nContactzhuhongmei@genomics.cn\n\nSupplementary InformationSupplementary data are available at Bioinformatics online

bioinformatics

Biogeography of the savanna-like vegetation in hot dry valleys in southwestern China with reference to their floristic origin and evolution

Savanna-like vegetation and dry thickets occur in hot dry valleys across southwestern China. Here, the flora and biogeography of these vegetations are studied. Native seed plants of 3,217 species from 1,038 genera in 163 families are recorded from the hot dry valleys in SW China. The biogeographical elements with a tropical distribution contribute 57.18%, but the ones with a temperate distribution contribute 36.45% of the total genera of the flora. This shows that the flora has proliferated by temperate elements via their evolution, although the flora occur in tropical habitats in the hot dry valleys. Floristic divergence across these hot dry valleys is obvious. The floras in the Yuanjiang (the upper reaches of the Red River) and the Nujiang (the upper reaches of the Salween River) valleys are dominated by tropical elements (77.26% and 74.49 of the total genera, respectively), but the flora of the Jinshajiang (the upper reaches of the Yangtze River) valley is composed of half tropical (47.27%) and half temperate (44.96%) genera. Regarding floristic similarities, the Jinshajiang shows the highest similarity to the Yuanjiang although these river valleys are located a great distance from each other. Our results could be well explained from the geological events since the Cenozoic, such as the uplift of Himalayas, the extrusion of Indochina, the river capture of the Jinshajiang separating from the Yuanjiang, and the northward movement of the Burma Plate. Further floristic comparison between the flora in hot dry valleys of SW China and southern Africa supports the consideration that the flora of savanna-like vegetations of SW China could have floristic affinity to African savannas over the course of its evolutionary history by the Indian Plate from southern Africa colliding with Eurasia in the Cenozoic.

ecology

Integrative pathway enrichment analysis of multivariate omics data

Multi-omics datasets quantify complementary aspects of molecular biology and thus pose challenges to data interpretation and hypothesis generation. ActivePathways is an integrative method that discovers significantly enriched pathways across multiple omics datasets using a statistical data fusion approach, rationalizes contributing evidence and highlights associated genes. We demonstrate its utility by analyzing coding and non-coding mutations from 2,583 whole cancer genomes, revealing frequently mutated hallmark pathways and a long tail of known and putative cancer driver genes. We also studied prognostic molecular pathways in breast cancer subtypes by integrating genomic and transcriptomic features of tumors and tumor-adjacent cells and found significant associations with immune response processes and anti-apoptotic signaling pathways. ActivePathways is a versatile method that improves systems-level understanding of cellular organization in health and disease through integration of multiple molecular datasets and pathway annotations.

bioinformatics

Limits to anatomical accuracy of diffusion tractography using modern approaches

Diffusion MRI fiber tractography is widely used to probe the structural connectivity of thebrain, with a range of applications in both clinical and basic neuroscience. Despite widespread use, tractography has well-known pitfalls that limits the anatomical accuracy of this technique. Numerous modern methods have been developed to address these shortcomings through advances in acquisition, modeling, and computation. To test whether these advances improve tractography accuracy, we organized the ISBI 2018 3D Validation of Tractography with Experimental MRI (3D-VoTEM) challenge. We made available three unique independent tractography validation datasets - a physical phantom and two ex vivo brain specimens - resulting in 176 distinct submissions from 9 research groups. By comparing results over a wide range of fiber complexities and algorithmic strategies, this challenge provides a more comprehensive assessment of tractographys inherent limitations than has been reported previously. The central results were consistent across all sub-challenges in that, despite advances in tractography methods, the anatomical accuracy of tractography has not dramatically improved in recent years. Taken together, our results independently confirm findings from decades of tractography validation studies, demonstrate inherent limitations in reconstructing white matter pathways using diffusion MRI data alone, and highlight the need for alternative or combinatorial strategies to accurately map the fiber pathways of the brain.

neuroscience

Development and Application of a High-Content Virion Display Human GPCR Array

G protein-coupled receptors (GPCRs) comprise the largest membrane protein family in humans and can respond to a wide variety of ligands and stimuli. Like other multi-pass membrane proteins, the biochemical properties of GPCRs are notoriously difficult to study because they must be embedded in lipid bilayers to maintain their native conformation and function. To enable an unbiased, high-throughput platform to profile biochemical activities of GPCRs in native conformation, we individually displayed 315 human non-odorant GPCRs (>85% coverage) in the envelope of human herpes simplex virus-1 and immobilized on glass to form a high-content Virion Display (VirD) array. Using this array, we found that 50% of the tested commercial anti-GPCR antibodies (mAbs) is ultra-specific, and that the vast majority of those VirD-GPCRs, which failed to be recognized by the commercial mAbs, could bind to their canonical ligands, indicating that they were folded correctly. Next, we used the VirD-GPCR arrays to examine binding specificity of two known peptide ligands and recovered expected interactions, as well as new off-target interactions, three of which were confirmed with real-time kinetics measurements. Finally, we explored the possibility of discovering novel pathogen targets by probing VirD-GPCR arrays with live group B Streptococcus (GBS), a common Gram-positive bacterium causing neonatal meningitis. Using cell invasion assays and a mouse model of hematogenous meningitis, we showed that inhibition of one of the five newly identified GPCRs, CysLTR1, greatly reduced GBS penetration in brain-derived endothelial cells and in mouse brains. Therefore, our work demonstrated that the VirD-GPCR array holds great potential for high-throughput, unbiased screening for small molecule drugs, affinity reagents, and deorphanization.

pharmacology and toxicology

Parkinson-associated SNCA enhancer variants revealed by open chromatin in mouse dopamine neurons

The progressive loss of midbrain (MB) dopaminergic (DA) neurons defines the motor features of Parkinson disease (PD) and modulation of risk by common variation in PD has been well established through GWAS. Anticipating that a fraction of PD-associated genetic variation mediates their effects within this neuronal population, we acquired open chromatin signatures of purified embryonic mouse MB DA neurons. Correlation with >2,300 putative enhancers assayed in mice reveals enrichment for MB cis-regulatory elements (CRE), data reinforced by transgenic analyses of six additional sequences in zebrafish and mice. One CRE, within intron 4 of the familial PD gene SNCA, directs reporter expression in catecholaminergic neurons of transgenic mice and zebrafish. Sequencing of this CRE in 986 PD patients and 992 controls reveals two common variants associated with elevated PD risk. To assess potential mechanisms of action, we screened >20,000 DNA interacting proteins and identify a subset whose binding is impacted by these enhancer variants. Additional genotyping across the SNCA locus identifies a single PD-associated haplotype, containing the minor alleles of both of the aforementioned PD-risk variants. Our work posits a model for how common variation at SNCA may modulate PD risk and highlights the value of cell context-dependent guided searches for functional non-coding variation.

genetics

Recovered and dead outcome patients caused by influenza A (H7N9) virus infection show different pro-inflammatory cytokine dynamics during disease progress and its application in real-time prognosis

The persistent circulation of influenza A(H7N9) virus within poultry markets and human society leads to sporadic epidemics of influenza infections. Severe pneumonia and acute respiratory distress syndrome (ARDS) caused by the virus lead to high morbidity and mortality rates in patients. Hyper induction of pro-inflammatory cytokines, which is known as \"cytokine storm\", is closely related to the process of viral infection. However, systemic analyses of H7N9 induced cytokine storm and its relationship with disease progress need further illuminated. In our study we collected 75 samples from 24 clinically confirmed H7N9-infected patients at different time points after hospitalization. Those samples were divided into three groups, which were mild, severe and fatal groups, according to disease severity and final outcome. Human cytokine antibody array was performed to demonstrate the dynamic profile of 80 cytokines and chemokines. By comparison among different prognosis groups and time series, we provide a more comprehensive insight into the hypercytokinemia caused by H7N9 influenza virus infection. Different dynamic changes of cytokines/chemokines were observed in H7N9 infected patients with different severity. Further, 33 cytokines or chemokines were found to be correlated with disease development and 11 of them were identified as potential therapeutic targets. Immuno-modulate the cytokine levels of IL-8, IL-10, BLC, MIP-3a, MCP-1, HGF, OPG, OPN, ENA-78, MDC and TGF-{beta} 3 are supposed to be beneficial in curing H7N9 infected patients. Apart from the identification of 35 independent predictors for H7N9 prognosis, we further established a real-time prediction model with multi-cytokine factors for the first time based on maximal relevance minimal redundancy method, and this model was proved to be powerful in predicting whether the H7N9 infection was severe or fatal. It exhibited promising application in prognosing the outcome of a H7N9 infected patients and thus help doctors take effective treatment strategies accordingly.

immunology

Cryo-EM structure of the benzodiazepine-sensitive α1β1γ2 heterotrimeric GABAA receptor in complex with GABA illuminates mechanism of receptor assembly and agonist binding

Fast inhibitory neurotransmission in the mammalian nervous system is largely mediated by GABAA receptors, chloride-selective members of the superfamily of pentameric Cys-loop receptors. Native GABAA receptors are heteromeric assemblies sensitive to many important drugs, from sedatives to anesthetics and anticonvulsive agents, with mutant forms of GABAA receptors implicated in multiple neurological diseases, including epilepsy. Despite the profound importance of heteromeric GABAA receptors in neuroscience and medicine, they have proven recalcitrant to structure determination. Here we present the structure of the triheteromeric 1{beta}1{gamma}2EM GABAA receptor in complex with GABA, determined by single particle cryo-EM at 3.1-3.8 [A] resolution, elucidating the molecular principles of receptor assembly and agonist binding. Remarkable N-linked glycosylation on the 1 subunit occludes the extracellular vestibule of the ion channel and is poised to modulate receptor assembly and perhaps ion channel gating. Our work provides a pathway to structural studies of heteromeric GABAA receptors and a framework for the rational design of novel therapeutic agents.

biophysics

Molecular Subtypes of Anaplastic Gliomas Identified with Somatic-Mutation and Pathway Based Gene Signature

PurposeAnaplastic gliomas constitute heterogeneous population with variable outcomes and no consensus on therapeutic approach. Molecular profiling may provide prognostication beyond clinical and pathologic factors and help guide treatment decisions.\n\nExperimental DesignThe Cancer Genome Atlas (TCGA) was utilized to derive a 39-gene low grade glioma-specific gene signature. Consensus clustering based on expression of the signature identified subgroups for 176 patients with anaplastic glioma from TCGA. Overall survival (OS) was analyzed for each subgroup. A total of 68 patients from Repository for Molecular Brain Neoplasia Data (REMBRANDT) were used as an independent validation dataset.\n\nResultsConsensus clustering separated the TCGA group into two distinct cohorts. The OS was significantly different between two subgroups, 20 vs. 67 months (p<0.001). On univariate analysis, the molecular subgroup, age, KPS, IDH1/2 mutation, 1p19q-co-deletion, chemotherapy, and use of both chemotherapy and radiation-therapy were significantly associated with OS. On multivariable analysis, the molecular subgroup remained significant with HR of 2.6 (p=.047, 95%CI [1.01-6.68]). In an independent validation with REMBRANDT, consensus clustering based on the signature successfully identified similarly poor prognostic subgroup with median survival of 14 months and concordance of expression patterns in 21 of the genes.\n\nConclusionExpression patterns of the 39 gene stratified anaplastic gliomas into two distinct subgroups with substantially different OS. This molecular prognostication was validated in an external dataset. Utilization of molecular subgroup, in addition to known prognostic factors may help define those requiring aggressive therapeutic intervention. Characteristic genes within the poor prognostic group may represent potential targets for therapeutic intensification.\n\nSource code and dataset used in this work is available for reviewers at: https://www.taehyunlab.org/ntripath\n\nImportance of the studyDespite the revolution of tailored therapy, anaplastic gliomas represent a category of tumors without clear treatment recommendations. While current prognostic factors help guide therapy recommendations, further refinement with the addition of molecular markers can help physicians with treatment recommendations. In this study, we developed a 39 gene prognostic gene signature by utilizing Pan-Cancer TCGA mutation profiles of over 5,000 patients across 19 different TCGA cancer types to stratify patients with anaplastic gliomas. We performed consensus clustering based on gene expression profiles of our 39-gene signature without any consideration of clinical factors or outcomes to TCGA anaplastic gliomas patients as well as an independent dataset and successfully identified two molecular subgroups with distinct clinical outcome. We found that subgroups have remarkably different survivals with a clear poor prognostic group. Furthermore, the poor prognostic group showed significant benefits for aggressive multimodality therapy, justifying the use of intensive therapy based on molecular stratification.

genomics

Database-integrated genome screening (DIGS): exploring genomes heuristically using sequence similarity search tools and a relational database.

A significant fraction of most genomes is comprised of DNA sequences that have been incompletely investigated. This genomic dark matter contains a wealth of useful biological information that can be recovered by systematically screening genomes in silico using sequence similarity search tools. Specialized computational tools are required to implement these screens efficiently. Here, we describe the database-integrated genome-screening (DIGS) tool: a computational framework for performing these investigations. To demonstrate, we screen mammalian genomes for endogenous viral elements (EVEs) derived from the Filoviridae, Parvoviridae, Circoviridae and Bornaviridae families, identifying numerous novel elements in addition to those that have been described previously. The DIGS tool provides a simple, robust framework for implementing a broad range of heuristic, sequence analysis-based explorations of genomic diversity.\n\nAvailabilityhttp://giffordlabcvr.github.io/DIGS-tool/\n\nContactrobert.gifford@glasgow.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Large-scale neuroimaging and genetic study reveals genetic architecture of brain white matter microstructure

Microstructural changes of white matter (WM) tracts are known to be associated with various neuropsychiatric disorders/diseases. Heritability of structural changes of WM tracts has been examined using diffusion tensor imaging (DTI) in family-based studies for different age groups. The availability of genetic and DTI data from recent large population-based studies offers opportunity to further improve our understanding of genetic contributions. Here, we analyzed the genetic architecture of WM tracts using DTI and single-nucleotide polymorphism (SNP) data of unrelated individuals in the UK Biobank (n [~] 8000). The DTI parameters were generated using the ENIGMA-DTI pipeline. We found that DTI parameters are substantially heritable on most WM tracts. We observed a highly polygenic or omnigenic architecture of genetic influence across the genome as well as the enrichment of SNPs in active chromatin regions. Our bivariate analyses showed strong genetic correlations for several pairs of WM tracts as well as pairs of DTI parameters. We performed voxel-based analysis to illustrate the pattern of genetic effects on selected parts of the tract-based spatial statistics skeleton. Comparing the estimates from the UK Biobank to those from small population-based studies, we illustrated that sufficiently large sample size is essential for genetic architecture discovery in imaging genetics. We confirmed this finding with a simulation study.

genetics

NMR Protein Structure Determination via Local Conformational Mapping

The ability of proteins to adopt multiple conformational states is essential to their function and elucidating the details of such diversity under physiological conditions has been a major challenge. Here we present a generalized method for mapping protein population landscapes by NMR spectroscopy. Experimental NOESY spectra are directly compared to a set of expectation spectra back-calculated across an arbitrary conformational space. Signal decomposition of the experimental spectrum then directly yields the relative populations of local conformational microstates. In this way, averaged descriptions of conformation can be eliminated. As the method quantitatively compares experimental and expectation spectra, it inherently delivers an R-factor expressing how well structural models explain the input data. We demonstrate that our method extracts sufficient information from a single 3D NOESY experiment to perform initial model building, refinement and validation, thus offering a complete de novo structure determination protocol.

biophysics

Membrane proteins with high N-glycosylation, high expression, and multiple interaction partners were preferred by mammalian viruses as receptors

Receptor mediated entry is the first step for viral infection. However, the relationship between viruses and receptors is still obscure. Here, by manually curating a high-quality database of 268 pairs of mammalian virus-host receptor interaction, which included 128 unique viral species or sub-species and 119 virus receptors, we found the viral receptors were structurally and functionally diverse, yet they had several common features when compared to other cell membrane proteins: more protein domains, higher level of N-glycosylation, higher ratio of self-interaction and more interaction partners, and higher expression in most tissues of the host. Additionally, the receptors used by the same virus tended to co-evolve. Further correlation analysis between viral receptors and the tissue and host specificity of the virus shows that the virus receptor similarity was a significant predictor for mammalian virus cross-species. This work could deepen our understanding towards the viral receptor selection and help evaluate the risk of viral zoonotic diseases.

microbiology

CRISPR-DT: designing gRNAs for the CRISPR-Cpf1 system with improved target efficiency and specificity

The CRISPR-Cpf1 system has been successfully applied in genome editing. However, target efficiency of the CRISPR-Cpf1 system varies among different gRNA sequences. We reanalyzed the published CRISPR-Cpf1 gRNAs data and found many sequence and structural features related to their target efficiency. Using machine learning technology, a SVM model was created to predict target efficiency for any given gRNAs. We have developed the first web service application, CRISPR-DT (CRISPR DNA Targeting), to help users design optimal gRNAs for the CRISPR-Cpf1 system by considering both target efficiency and specificity. CRISPR-DT is available at http://bioinfolab.miamioh.edu/CRISPR-DT.

bioinformatics

A Novel QconCAT-Based Proteomics Method for Determining Allele-Specific Protein Expression (ASPE): a New Approach to Identify Cis-acting Genetic Variants

Measuring allele-specific expression (ASE) is a powerful approach for identifying cis-regulatory genetic variants. Here we developed a novel targeted proteomics method for quantification of allele-specific protein expression (ASPE) based on scheduled high resolution multiple reaction monitoring (sMRM-HR) with a heavy stable isotope-labeled quantitative concatamer (QconCAT) internal protein standard. This strategy was applied to the determination of the ASPE of UGT2B15 in human livers using the common UGT2B15 nonsynonymous variant rs1902023 (i.e. Y85D) as the marker to differentiate expressions from the two alleles. The QconCAT standard contains both the wild type tryptic peptide and the Y85D mutant peptide at a ratio of 1:1 to ensure accurate measurement of the ASPE of UGT2B15. The results from 18 UGT2B15 Y85D heterozygotes revealed that the ratios between wild type Y allele and mutant D allele varied from 0.60 to 1.46, indicating the presence of cis-regulatory variants. In addition, we observed no significant correlations between the ASPE and mRNA ASE of UGT2B15, suggesting the involvement of different cis-acting variants in regulating the transcription and translation processes of the gene. This novel ASPE approach provides a powerful tool for capturing cis-genetic variants involved in post-transcription processes, an important yet understudied area of research.

molecular biology