Search bioRxivSearch

Biology subjects

Zhang, S.

Publications and source records attributed to Zhang, S..

At least 55 records · Page 3Linked to original sources

Testing an Optimally Weighted Combination of Common and/or Rare Variants with Multiple Traits

Joint analysis of multiple traits has recently become popular since it can increase statistical power to detect genetic variants and there is increasing evidence showing that pleiotropy is a widespread phenomenon in complex diseases. Currently, most of existing methods test the association between multiple traits and a single common variant. However, the variant-by-variant methods for common variant association studies may not be optimal for rare variant association studies due to the allelic heterogeneity as well as the extreme rarity of individual variants. In this article, we developed a statistical method by testing an optimally weighted combination of variants with multiple traits (TOWmuT) to test the association between multiple traits and a weighted combination of variants (rare and/or common) in a genomic region. TOWmuT is robust to the directions of effects of causal variants and is applicable to different types of traits. Using extensive simulation studies, we compared the performance of TOWmuT with the following five existing methods: gene association with multiple traits (GAMuT), multiple sequence kernel association test (MSKAT), adaptive weighting reverse regression (AWRR), single-TOW, and MANOVA. Our results showed that, in all of the simulation scenarios, TOWmuT has correct type I error rates and is consistently more powerful than the other five tests. We also illustrated the usefulness of TOWmuT by analyzing a whole-genome genotyping data from a lung function study.

bioinformatics

A nuclear hormone receptor and lipid metabolism axis are required for the maintenance and regeneration of reproductive organs

Understanding how stem cells and their progeny maintain and regenerate reproductive organs is of fundamental importance. The freshwater planarian Schmidtea mediterranea provides an attractive system to study these processes because its hermaphroditic reproductive system (RS) arises post-embryonically and when lost can be fully and functionally regenerated from the proliferation and regulation of experimentally accessible stem and progenitor cells. By controlling the function of a nuclear hormone receptor gene (nhr-1), we established conditions in which to study the formation, maintenance and regeneration of both germline and somatic tissues of the planarian RS. We found that nhr-1(RNAi) not only resulted in the gradual degeneration and complete loss of the adult hermaphroditic RS, but also in the significant downregulation of a large cohort of genes associated with lipid metabolism. One of these, Smed-acs-1, a homologue of Acyl-CoA synthetase, was indispensable for the development, maintenance and regeneration of the RS, but not for the homeostasis or regeneration of other somatic tissues. Remarkably, supplementing nhr-1(RNAi) animals with either bacterial Acyl-CoA synthetase or the lipid metabolite Acetyl-CoA rescued the phenotype restoring the maintenance and function of the hermaphroditic RS. Our findings uncovered a likely evolutionarily conserved role for nuclear hormone receptors and lipid metabolism in the regulation of stem and progenitor cells required for the long-term maintenance and regeneration of animal reproductive organs, tissues and cells.

developmental biology

Identification of novel mutations associated with clofazimine resistance in Mycobacterium abscessus

Mycobacterium abscessus (Mab) is a major non-tuberculous mycobacterial (NTM) pathogen responsible for about 80% of all pulmonary infections caused by rapidly growing mycobacteria. Clofazimine is an effective drug active against Mab and shows synergistic activity when given with amikacin, but the mechanism of resistance to clofazimine in Mab is unknown.\n\nObjectiveTo investigate the molecular basis of clofazimine resistance in Mab.\n\nMethodsWe isolated 29 Mab mutants resistant to clofazimine, and subjected them to whole genome sequencing and Sanger sequencing to identify possible mutations associated with clofazimine resistance.\n\nResultsMutations in MAB_2299c gene which encodes possible transcriptional regulatory protein were identified in 23 of the 29 clofazimine-resistant mutants. In addition, 6 mutations in MAB_1483 were found in 21 of the 29 mutants, and one mutation in MAB_0540 was found in 16 of the 29 mutants. Mutations in MAB_0416c, MAB_4099c, MAB_2613, MAB_0409, MAB_1426 were also associated with clofazimine resistance in less frequency. Two identical mutations which are likely to be polymorphisms unrelated to clofazimine resistance were found in MAB_4605c and MAB_4323 in 13 mutants.\n\nConclusionMutations in MAB_2299c, MAB_1483, and MAB_0540 are the major mechanisms of clofazimine resistance in Mab. Future studies are needed to address the role of the identified mutations in clofazimine resistance in Mab, and our findings have implications for developing a rapid molecular test for detecting clofazimine resistance in this organism.

microbiology

Learning common and specific patterns from data of multiple interrelated biological scenarios with matrix factorization

High-throughput biological technologies (e.g., ChIP-seq, RNA-seq and single-cell RNA-seq) rapidly accelerate the accumulation of genome-wide omics data in diverse interrelated biological scenarios (e.g., cells, tissues and conditions). Data dimension reduction and differential analysis are two common paradigms for exploring and analyzing such data. However, they are typically used in a separate or/and sequential manner. In this study, we propose a flexible non-negative matrix factorization framework CSMF to combine them into one paradigm to simultaneously reveal common and specific patterns from data generated under interrelated biological scenarios. We demonstrate the effectiveness of CSMF with four applications including pairwise ChIP-seq data describing the chromatin modification map on protein-DNA interactions between K562 and Huvec cell lines; pairwise RNA-seq data representing the expression profiles of two cancers (breast invasive carcinoma and uterine corpus endometrial carcinoma); RNA-seq data of three breast cancer subtypes; and single-cell sequencing data of human embryonic stem cells and differentiated cells at six time points. Extensive analysis yields novel insights into hidden combinatorial patterns embedded in these interrelated multi-modal data. Results demonstrate that CSMF is a powerful tool to uncover common and specific patterns with significant biological implications from data of interrelated biological scenarios.

bioinformatics

The proteome of the malaria plastid organelle, a key anti-parasitic target

Malaria parasites (Plasmodium spp.) and related apicomplexan pathogens contain a non-photosynthetic plastid called the apicoplast. Derived from an unusual secondary eukaryote-eukaryote endosymbiosis, the apicoplast is a fascinating organelle whose function and biogenesis rely on a complex amalgamation of bacterial and algal pathways. Because these pathways are distinct from the human host, the apicoplast is an excellent source of novel antimalarial targets. Despite its biomedical importance and evolutionary significance, the absence of a reliable apicoplast proteome has limited most studies to the handful of pathways identified by homology to bacteria or primary chloroplasts, precluding our ability to study the most novel apicoplast pathways. Here we combine proximity biotinylation-based proteomics (BioID) and a new machine learning algorithm to generate a high-confidence apicoplast proteome consisting of 346 proteins. Critically, the high accuracy of this proteome significantly outperforms previous prediction-based methods and extends beyond other BioID studies of unique parasite compartments. Half of identified proteins have unknown function, and 77% are predicted to be important for normal blood-stage growth. We validate the apicoplast localization of a subset of novel proteins and show that an ATP-binding cassette protein ABCF1 is essential for blood-stage survival and plays a previously unknown role in apicoplast biogenesis. These findings indicate critical organellar functions for newly discovered apicoplast proteins. The apicoplast proteome will be an important resource for elucidating unique pathways derived from secondary endosymbiosis and prioritizing antimalarial drug targets.

microbiology

DeepHINT: Understanding HIV-1 integration via deep learning with attention

MotivationHuman immunodeficiency virus type 1 (HIV-1) genome integration is closely related to clinical latency and viral rebound. In addition to human DNA sequences that directly interact with the integration machinery, the selection of HIV integration sites has also been shown to depend on the heterogeneous genomic context around a large region, which greatly hinders the prediction and mechanistic studies of HIV integration.\n\nResultsWe have developed an attention-based deep learning framework, named DeepHINT, to simultaneously provide accurate prediction of HIV integration sites and mechanistic explanations of the detected sites. Extensive tests on a high-density HIV integration site dataset showed that DeepHINT can outperform conventional modeling strategies by automatically learning the genomic context of HIV integration solely from primary DNA sequence information. Systematic analyses on diverse known factors of HIV integration further validated the biological relevance of the prediction result. More importantly, in-depth analyses of the attention values output by DeepHINT revealed intriguing mechanistic implications in the selection of HIV integration sites, including potential roles of several basic helix-loop-helix (bHLH) transcription factors and zinc-finger proteins. These results established DeepHINT as an effective and explainable deep learning framework for the prediction and mechanistic study of HIV integration.\n\nAvailabilityDeepHINT is available as an open-source software and can be downloaded from https://github.com/nonnerdling/DeepHINT\n\nContactlzhang20@mail.tsinghua.edu.cn and zengjy321@tsinghua.edu.cn

bioinformatics

Gut bacterial metabolite Urolithin A (UA) mitigates Ca2+ 1 entry inT cells by regulating miR-10a-5p

The gut microbiota influences several biological functions including immune response. Inflammatory bowel disease is favourably influenced by consumption of several dietary natural plant products such as pomegranate, walnuts and berries containing polyphenolic compounds such as ellagitannins and ellagic acid. The gut microbiota metabolises ellagic acid leading to formation of bioactive urolithins A, B, C and D. Urolithin A (UA) is the most active and effective gut metabolite and acts as a potent anti-inflammatory and anti-oxidant agent. However, how gut metabolite UA affects the function of immune cells remained incompletely understood. T cell proliferation is stimulated by store operated Ca2+ entry (SOCE) resulting from stimulation of Orai1 by STIM1/STIM2. We show here that treatment of murine CD4+ T cells with UA (10 {micro}M, 3 days) significantly blunted SOCE in CD4+ T cells, an effect paralleled by significant downregulation of Orai1 and STIM1/2 transcript levels and protein abundance. UA treatment further increased miR-10a-5p abundance in CD4+ T cells in a dose dependent fashion. Overexpression of miR-10a-5p significantly decreased STIM1/2 and Orai1 mRNA and protein levels as well as SOCE in CD4+ T cells. UA further decreased CD4+ T cell proliferation. Thus, bacterial metabolite UA up-regulates miR-10a-5p thus interfering with Orai1/STIM1/STIM2 expression, store operated Ca2+ entry and proliferation of murine CD4+ T cells.

immunology

Mutations in efflux pump Rv1258c (Tap) cause resistance to pyrazinamide and other drugs in M. tuberculosis

Although drug resistance in M. tuberculosis is mainly caused by mutations in drug activating enzymes or drug targets, there is increasing interest in possible role of efflux in causing drug resistance. Previously, efflux genes are shown upregulated upon drug exposure or implicated in drug resistance in overexpression studies, but the role of mutations in efflux pumps identified in clinical isolates in causing drug resistance is unknown. Here we investigated the role of mutations in efflux pump Rv1258c (Tap) from clinical isolates in causing drug resistance in M. tuberculosis by constructing point mutations V219A, S292L in Rv1258c in the chromosome of M. tuberculosis and assessed drug susceptibility of the constructed mutants. Interestingly, V219A, S292L point mutations caused clinically relevant drug resistance to pyrazinamide (PZA), isoniazid (INH), and streptomycin (SM), but not to other drugs in M. tuberculosis. While V219A point mutation conferred a low level resistance, the S292L mutation caused a higher level of resistance. Efflux inhibitor piperine inhibited INH and PZA resistance in the S292L mutant but not in the V219A mutant. S292L mutant had higher efflux activity for pyrazinoic acid (the active form of PZA) than the parent strain. We conclude that point mutations in the efflux pump Rv1258c in clinical isolates can confer clinically relevant drug resistance including PZA and could explain some previously unaccounted drug resistance in clinical strains. Future studies need to take efflux mutations into consideration for improved detection of drug resistance in M. tuberculosis and address their role in affecting treatment outcome in vivo.

microbiology

Identification of novel mutations in LprG (rv1411c), rv0521, rv3630, rv0010c, ppsC, cyp128 associated with pyrazinoic acid/pyrazinamide resistance in Mycobacterium tuberculosis

There is currently considerable interest in understanding the mechanisms of action of pyrazinamide (PZA), a critical frontline tuberculosis (TB) drug that plays a unique role in shortening TB therapy due to its unique activity against Mycobacterium tuberculosis persisters that are not killed by other TB drugs.1,2 Despite the importance of PZA in the treatment of both drug susceptible and drug-resistant TB and its simple structure, its mechanisms of action are complex and are not well understood.1,2 PZA is a prodrug that is converted to the active form pyrazinoic acid (POA) by nicotinamidase/pyrazinamidase (PZase) encoded by the pncA gene,3 whose mutation is the most common mechanism of PZA resistance in M. tuberculosis.3-5 However, some low level PZA-resistant strains (MIC=200-300 g/ml, pH6.0) do not have mutations in the pncA gene.5,6 Recent studies have identified rpsA, which encodes the ribosomal protein S1 involved in both translation and trans-translation process, as a target of PZA,7 where its mutations are associated with PZA resistance from clinical isolates. In addition, mutations in panD encoding aspartate decarboxylase were identified as a new mechanism of PZA resistance from in vitro mutants resistant to PZA and the PanD protein was found to be another target of PZA.8,9 panD mutations were initially found in mutants resistant to PZA 8 and then in mutants resistant to POA. 9,10 It is worth noting that previously POA-resistant mutants could not be isolated at acid pH which is required for higher activity of PZA against M. tuberculosis. However, we were able to successfully isolate POA-resistant mutants with high POA concentrations at close to neutral pH (pH 6.8),9 which led to discovery of new genes involved in POA and PZA resistance. For example, clpC1, which was also isolated from mutants resistant to PZA,11 was identified in mutants resistant to POA.12

microbiology

Proteasome Substrate Capture and Gate Activation by Mycobacterium tuberculosis PafE

In all domains of life, proteasomes are gated chambered proteases that require opening by activators in order to facilitate protein degradation. Twelve proteasome accessory factor E (PafE) monomers assemble into a single, dodecameric ring to promote proteolysis that is required for the full virulence of the human bacterial pathogen Mycobacterium tuberculosis. While the best characterized proteasome activators use ATP to deliver proteins into a proteasome, PafE does not require ATP. In order to understand the mechanism of PafE-mediated protein targeting and proteasome activation, we studied the interactions of PafE with native substrates, including a newly identified proteasome substrate, Rv3213c, and with proteasome core particles. We characterized the function of a highly conserved feature conserved in bacterial proteasome activator proteins: a glycine-glutamine-tyrosine-leucine or \"GQYL\" motif at their carboxyl-termini that is essential to stimulate proteolysis. Using cryo-electron microscopy, we found that the GQYL motif of PafE interacts with specific residues in the -subunits of the proteasome core particle to trigger gate opening and degradation. Finally, we found that PafE rings have 40-[A] openings lined with hydrophobic residues that form a chamber for capturing substrates prior to the onset of degradation. This result suggests PafE has a previously unrecognized chaperone activity. Collectively, our data provide new insights on the mechanistic understanding of ATP-independent proteasome degradation in bacteria.

biochemistry

Identifying and exploiting trait-relevant tissues with multiple functional annotations in genome-wide association studies

Genome-wide association studies (GWASs) have identified many disease associated loci, the majority of which have unknown biological functions. Understanding the mechanism underlying trait associations requires identifying trait-relevant tissues and investigating associations in a trait-specific fashion. Here, we extend the widely used linear mixed model to incorporate multiple SNP functional annotations from omics studies with GWAS summary statistics to facilitate the identification of trait-relevant tissues, with which to further construct powerful association tests. Specifically, we rely on a generalized estimating equation based algorithm for parameter inference, a mixture modeling framework for trait-tissue relevance classification, and a weighted sequence kernel association test constructed based on the identified trait-relevant tissues for powerful association analysis. We refer to our analytic procedure as the Scalable Multiple Annotation integration for trait-Relevant Tissue identification and usage (SMART). With extensive simulations, we show how our method can make use of multiple complementary annotations to improve the accuracy for identifying trait-relevant tissues. In addition, our procedure allows us to make use of the inferred trait-relevant tissues, for the first time, to construct more powerful SNP set tests. We apply our method for an in-depth analysis of 43 traits from 28 GWASs using tissue-specific annotations in 105 tissues derived from ENCODE and Roadmap. Our results reveal new trait-tissue relevance, pinpoint important annotations that are informative of trait-tissue relationship, and illustrate how we can use the inferred trait-relevant tissues to construct more powerful association tests in the Wellcome trust case control consortium study.\n\nAuthor SummaryIdentifying trait-relevant tissues is an important step towards understanding disease etiology. Computational methods have been recently developed to integrate SNP functional annotations generated from omics studies to genome-wide association studies (GWASs) to infer trait-relevant tissues. However, two important questions remain to be answered. First, with the increasing number and types of functional annotations nowadays, how do we integrate multiple annotations jointly into GWASs in a trait-specific fashion to take advantage of the complementary information contained in these annotations to optimize the performance of trait-relevant tissue inference? Second, what to do with the inferred trait-relevant tissues? Here, we develop a new statistical method and software to make progress on both fronts. For the first question, we extend the commonly used linear mixed model, with new algorithms and inference strategies, to incorporate multiple annotations in a trait-specific fashion to improve trait-relevant tissue inference accuracy. For the second question, we rely on the close relationship between our proposed method and the widely-used sequence kernel association test, and use the inferred trait-relevant tissues, for the first time, to construct more powerful association tests. We illustrate the benefits of our method through extensive simulations and applications to a wide range of real data sets.

bioinformatics

Comparison of computational methods for imputing single-cell RNA-sequencing data

Single-cell RNA-sequencing (scRNA-seq) is a recent breakthrough technology, which paves the way for measuring RNA levels at single cell resolution to study precise biological functions. One of the main challenges when analyzing scRNA-seq data is the presence of zeros or dropout events, which may mislead downstream analyses. To compensate the dropout effect, several methods have been developed to impute gene expression since the first Bayesian-based method being proposed in 2016. However, these methods have shown very diverse characteristics in terms of model hypothesis and imputation performance. Thus, large-scale comparison and evaluation of these methods is urgently needed now. To this end, we compared eight imputation methods, evaluated their power in recovering original real data, and performed broad analyses to explore their effects on clustering cell types, detecting differentially expressed genes, and reconstructing lineage trajectories in the context of both simulated and real data. Simulated datasets and case studies highlight that there are no one method performs the best in all the situations. Some defects of these methods such as scalability, robustness and unavailability in some situations need to be addressed in future studies.

bioinformatics

Clustering enzymes using E.coli inner cell membrane as scaffold in metabolic pathway

Clustering enzymes in the same metabolism pathway is a natural strategy to enhance the productivity. Several systems have been designed to artificially cluster desired enzymes in the cell, such as synthetic protein scaffold and nucleic acid scaffold. However, these scaffolds require complicated construction process and have limited slots for target enzymes. Following this direction, we designed a scaffold system based on natural cell membrane. Target enzymes (FabZ, FabG, FabI and TesA in fatty acid synthesis II pathway) are anchored on the E.coli inner membrane, showing the enhanced metabolism flux without the requirement of the further artificial interactions to force the clustering. Furthermore, anchoring the enzymes on the membrane enhances the products exportation, which further increases the productivity. Together, the proposed system has potential applications in producing valuable biomaterials.

synthetic biology

CLAN: the CrossLinked reads ANalysis tool

The crosslinked RNA sequencing technology ligates interacting RNA strands followed by next-generation sequencing. Mapping of the resulting duplex reads allows for functional inference of the corresponding intramolecular/intermolecular RNA-RNA interactions. However, duplex read mapping remains computationally challenging, and the existing best-performing software fails to map a significant portion of the duplex reads. To address this challenge, we develop a novel algorithm for duplex read mapping, called CrossLinked reads ANalysis tool (CLAN). CLAN demonstrates drastically improved sensitivity and high alignment accuracy when applied to real crosslinked RNA sequencing data. CLAN is implemented in GNU C++, and is freely available from http://sourceforge.net/projects/clan-mapping.

bioinformatics

The Control of Tonic Pain by Active Relief Learning

Tonic pain after injury characterises a behavioural state that prioritises recovery. Although generally suppressing cognition and attention, tonic pain needs to allow effective relief learning so that the cause of the pain can be reduced if possible. Here, we describe a central learning circuit that supports learning of relief and concurrently suppresses the level of ongoing pain. We used computational modelling of behavioural, physiological and neuroimaging data in two experiments in which subjects learned to terminate tonic pain in static and dynamic escape-learning paradigms. In both studies, we show that active relief-seeking involves a reinforcement learning process manifest by error signals observed in the dorsal putamen. Critically, this system also uses an uncertainty ( associability) signal detected in pregenual anterior cingulate cortex that both controls the relief learning rate, and endogenously and parametrically modulates the level of tonic pain. The results define a self-organising learning circuit that allows reduction of ongoing pain when learning about potential relief.

neuroscience

Selective Filopodia Adhesion Ensures Robust Cell Matching in the Drosophila Heart

The ability to form specific cell-cell connections within complex cellular environments is critical for multicellular organisms. However, the underlying mechanisms of cell matching that instruct these connections remain elusive. Here, we explore the dynamic regulation of matching processes utilizing Drosophila cardiogenesis. During embryonic heart formation, cardioblasts (CBs) form precise contacts with their partners after long-range migration. We find that CB matching is highly robust at the boundaries between distinct CB subtypes. Filopodia in these CB subtypes have different binding affinities. We identify the adhesion molecules Fasciclin III (Fas3) and Ten-m as having complementary differential expression in CBs. Altering Fas3 expression influences the CB filopodia selective binding activities and CB matching. In contrast to single knockouts, loss of both Fas3 and Ten-m dramatically impairs CB alignment. We propose that differential expression of adhesion molecules mediates selective filopodia binding, and these molecules work in concert to instruct precise and robust cell matching.

developmental biology

Newly identified relatives of botulinum neurotoxins shed light on their molecular evolution

The evolution of bacterial toxins is a central question to understanding the origins of human pathogens and infectious disease. Through genomic data mining, we traced the evolution of the deadliest known toxin family, clostridial neurotoxins, comprised of tetanus and botulinum neurotoxins (BoNT). We identified numerous uncharacterized lineages of BoNT-related genes in environmental species outside of Clostridium, revealing insights into their molecular ancestry. Phylogenetic analysis pinpointed a sister lineage of BoNT-like toxins in the gram-negative organism, Chryseobacterium piperi, that exhibit distant homology at the sequence level but preserve overall domain architecture. Resequencing and assembly of the C. piperi genome confirmed the presence of BoNT-like proteins encoded within two toxin-rich gene clusters. A C. piperi BoNT-like protein was validated as a novel toxin that induced necrotic cell death in human kidney cells. Mutagenesis of the putative active site abolished toxicity and indicated a zinc metalloprotease-dependent mechanism. The C. piperi toxin did not cleave common SNARE substrates of BoNTs, indicating that BoNTs have diverged from related families in substrate specificity. The new lineages of BoNT-like toxins identified by computational methods represent evolutionary missing links, and suggest an origin of clostridial neurotoxins from ancestral toxins present in environmental bacteria.\n\nSignificance statementThe origins of bacterial toxins that cause human disease is a key question in our understanding of pathogen evolution. To explore this question, we searched genomes for evolutionary relatives of the deadliest biological toxins known to science, botulinum neurotoxins. Genomic and phylogenetic analysis revealed a group of toxins in the Chryseobacterium piperi genome that are a sister lineage to botulinum toxins. Genome sequencing of this organism confirmed the presence of toxin-rich gene clusters, and a predicted C. piperi toxin was shown to induce necrotic cell death in human cells. These newly predicted toxins are missing links in our understanding of botulinum neurotoxin evolution, revealing its origins from an ancestral family of toxins that may be widespread in the environment.

bioinformatics

Single Molecule Sequencing of Cell-free DNA from Maternal Plasma for Noninvasive Trisomy Detection

The demand of non-invasive prenatal testing for autosomal aneuploidy using cell-free fetal DNA (cffDNA) in maternal plasma is a highly sought-after diagnostic, with a rapidly growing market. Current approaches developed by next generation sequencing (NGS) need PCR amplifcation during sample preparation, which results in amplification bias in GC-rich areas of the human genome. With these approaches, the minimum fetal fraction in maternal plasma is 4% for the small differences in circulating cfDNA between trisomic and disomic pregnancies to be detectable. In this paper, we performed single molecule sequencing of cell-free DNA from maternal plasma for noninvasive trisomy 13, 18 and 21 detections using the GenoCare platform. We found that single molecule sequencing is sensitive enough to detect these chromosome abnormalities when the fetal DNA fraction is as low as 2%. Compared to the Hiseq2500 platform, no significant GC bias was observed. The improved sensitivity and unbiased GC readout make GenoCare a promising platform for autosomal aneuploidy detections, even in the very early stage of pregnancy.

genomics