Search bioRxivSearch

Biology subjects

Zhao, X.

Publications and source records attributed to Zhao, X..

35 records · Page 2Linked to original sources

Targeted enrichment outperforms other enrichment techniques and enables more multi-species RNA-Seq analyses

Enrichment methodologies enable analysis of minor members in multi-species transcriptomic analyses. We compared standard enrichment of bacterial and eukaryotic mRNA to targeted enrichment with Agilent SureSelect (AgSS) capture for Brugia malayi, Aspergillus fumigatus, and the Wolbachia endosymbiont of B. malayi (wBm). Without introducing significant systematic bias, the AgSS quantitatively enriched samples, resulting in more reads mapping to the target organism. The AgSS-enriched libraries consistently had a positive linear correlation with its unenriched counterpart (r2=0.559-0.867). Up to a 2,242-fold enrichment of RNA from the target organism was obtained following a power law (r2=0.90), with the greatest fold enrichment achieved in samples with the largest ratio difference between the major and minor members. While using a single total library for prokaryote and eukaryote in a single sample could be beneficial for samples where RNA is limiting, we observed a decrease in reads mapping to protein coding genes and an increase of multi-mapping reads to rRNAs in AgSS enrichments from eukaryotic total RNA libraries as opposed to eukaryotic poly(A)-enriched libraries. Our results support a recommendation of using Agilent SureSelect targeted enrichment on poly(A)-enriched libraries for eukaryotic captures and total RNA libraries for prokaryotic captures to increase the robustness of multi-species transcriptomic studies.

genomics

Structural and functional influences of urban and rural childhoods on the medial prefrontal cortex

Global increases in urbanization have brought dramatic economic, environmental and social changes. However, less is understood about how these may influence disease-related brain mechanisms underlying epidemiological observations that urban birth and childhoods may increase the risk for neuropsychiatric disorders, including increased social stress and depression. In a genetically homogeneous Han Chinese adult population with divergent urban and rural birth and childhoods, we examined the structural and functional MRI neural correlates of childhood urbanicity, focusing on behavioral traits responding to social status threats, and polygenic risk for depression. Subjects with divergent rural and urban childhoods were similar in adult socioeconomic status and were genetically homogeneous. Urban childhoods, however, were associated with higher trait anxiety-depression. On structural MRI, urban childhoods were associated with relatively reduced medial prefrontal gray matter volumes. Functional medial prefrontal engagement under social status threat during working memory correlated with trait anxiety-depression in subjects with urban childhoods, to a significantly greater extent than in their rural counterparts, implicating an exaggerated physiological response to the threat context. Stress-associated medial prefrontal engagement also interacted with polygenic risk for depression, significantly predicting a differential response in individuals with urban but not rural childhoods. Developmental urbanicity thus differentially influenced medial prefrontal structure and function, at least in part through mechanisms associated with the neural processing of social status threat, trait anxiety, and genetic risk for depression, which may be factors in the association of urbanicity with adult psychopathology.\n\nSignificance StatementUrban living has been associated with social inequalities and stress. However, less is understood about the neural underpinnings by which these stressors affect disease risk, and in particular, genetic risk for depression. Leveraging urbanization in China, we studied adults with diverse urban and rural upbringings, who were genetically homogeneous and with similar current socioeconomic status, to isolate the effects of childhood urbanicity. At medial prefrontal cortex, a region critical for processing emotional stressors and social status, genetic risk for depression resulted in more deleterious function under stress in individuals with urban, but not rural childhoods. This implicates medial prefrontal cortexs critical role in brain development, integrating genetic mechanisms of stress and depression with the childhood environment.

neuroscience

A compromised gsdf signaling leads to gamatogenesis confusion and subfertility in medaka

Summary statementGsdf signals trigger the gamatogenesis, alter the somatic expression of Fsh/Lh receptors and brain type aromatase in medaka brain and gonad.\n\nAbstractGonadal soma-derived factor (gsdf) and anti-Mullerian hormone (amh) are somatic male determinants in several species of teleosts, although the mechanisms by which they trigger the indifferent germ cells into the male pathway remain unknown. This study aimed to decipher the roles of gsdf/amh in directing the sexual fate of germ cells using medaka as a model. Transgenic lines (TgcryG) that restrictively and persistently express a Gsdf-Gfp fusion protein in the lens and the hypothalamus-pituitary-gonad (HPG) axis, were generated under the control of a mouse {gamma}F-crystallin promoter. A high frequency (44.4%) of XX male sex reversals was obtained in TgcryG lines, indicating that signals of gsdf-expressing cells in HPG were enough for the spermatogenesis activation in the genetic females. Furthermore, all TgcryG XY individuals with endogenous gsdf depletion (named Sissy) displayed intersex (100%) with enlarged ovotestis in contrast to a giant ovary developed in XY gsdf deficiency. The heterogeneous expression of gsdf led to the confusion of gamatogenesis and ovotestis development, similar to some hotei (amhr2) mutants, suggests that the signaling balance of gsdf/amh is essential for proper gamatogenesis, maintaining sex steroid production and gonadotropin secretion, which are evolutionarily conserved across phyla.

evolutionary biology

NeuroField: Theory and Simulation of Multiscale Neural Field Dynamics

A user ready, portable, documented software package, NFTsim, is presented to facilitate numerical simulations of a wide range of brain systems using continuum neural field modeling. NFTsim enables users to simulate key aspects of brain activity at multiple scales. At the microscopic scale, it incorporates characteristics of local interactions between cells, neurotransmitter effects, synaptodendritic delays and feedbacks. At the mesoscopic scale, it incorporates information about medium to large scale axonal ranges of fibers, which are essential to model dissipative wave transmission and to produce synchronous oscillations and associated cross-correlation patterns as observed in local field potential recordings of active tissue. At the scale of the whole brain, NFTsim allows for the inclusion of long range pathways, such as thalamocortical projections, when generating macroscopic activity fields. The multiscale nature of the neural activity produced by NFTsim has the potential to enable the modeling of resulting quantities measurable via various neuroimaging techniques. In this work, we give a comprehensive description of the design and implementation of the software. Due to its modularity and flexibility, NFTsim enables the systematic study of an unlimited number of neural systems with multiple neural populations under a unified framework and allows for direct comparison with analytic and experimental predictions. The code is written in C++ and bundled with Matlab routines for a rapid quantitative analysis and visualization of the outputs. The output of NFTsim is stored in plain text file enabling users to select from a broad range of tools for offline analysis. This software enables a wide and convenient use of powerful physiologically-based neural field approaches to brain modeling. NFTsim is distributed under the Apache 2.0 license.

neuroscience

A Comprehensive Evaluation of the Genetic Architecture of Sudden Cardiac Arrest

BackgroundSudden cardiac arrest (SCA) accounts for 10% of adult mortality in Western populations. While several risk factors are observationally associated with SCA, the genetic architecture of SCA in the general population remains unknown. Furthermore, understanding which risk factors are causal may help target prevention strategies.\n\nMethodsWe carried out a large genome-wide association study (GWAS) for SCA (n=3,939 cases, 25,989 non-cases) to examine common variation genome-wide and in candidate arrhythmia genes. We also exploited Mendelian randomization methods using cross-trait multi-variant genetic risk score associations (GRSA) to assess causal relationships of 18 risk factors with SCA.\n\nResultsNo variants were associated with SCA at genome-wide significance, nor were common variants in candidate arrhythmia genes associated with SCA at nominal significance. Using cross-trait GRSA, we established genetic correlation between SCA and (1) coronary artery disease (CAD) and traditional CAD risk factors (blood pressure, lipids, and diabetes), (2) height and BMI, and (3) electrical instability traits (QT and atrial fibrillation), suggesting etiologic roles for these traits in SCA risk.\n\nConclusionsOur findings show that a comprehensive approach to the genetic architecture of SCA can shed light on the determinants of a complex life-threatening condition with multiple influencing factors in the general population. The results of this genetic analysis, both positive and negative findings, have implications for evaluating the genetic architecture of patients with a family history of SCA, and for efforts to prevent SCA in highrisk populations and the general community.

genetics

Deep-RBPPred: Predicting RNA binding proteins in the proteome scale based on deep learning

RNA binding protein (RBP) plays an important role in cell processes. Identifying RBPs by computation and experiment are both essential. Recently, RBPPred is proposed in our group to predict RBP with a high performance. However, RBPPred is too slow for that it will generate PSSM matrix as its feature. Herein, we develop a deep learning model called Deep-RBPPred. The model has three advantages comparing to previous models. 1. Deep-RBPPred only needs few physicochemical properties. 2. Deep-RBPPred runs much faster. 3. Deep-RBPPred has a good generalization ability. In the meantime, the performance is still as good as the stats-of-the-art method. In the testing in A. thaliana, S. cerevisiae and H. sapiens proteomics, MCC (AUC) are 0.6077 (0.9421), 0.573 (0.9034) and 0.8141(0.9515) respectively when the score cutoff is set to 0.5. In the verifying in Gerstberger-1538, the SN of our model is 90.38%. The running times are 9s, 7s, 8s and 10s, respectively, for H.sapiens, A.thaliana, S.cerevisiae and Gerstberger-1538 when it is tested in GPU. Deep-RBPPred forecasts 94.65% of 299 new RBP and about 8% higher sensitivity than RBPPred. We also apply deep-RBPPred in 19 eukaryotes proteomics and 11 bacteria proteomics downloaded from Uniprot. The result shows that rate of RBPs in eukaryotes proteome are much higher than bacteria proteome. Testing in 6 proteomics shows the many RBPs may be still undiscovered so far.

bioinformatics

A fine-tuned vector-parasite dialogue in tsetse’s cardia determines peritrophic matrix integrity and trypanosome transmission success

Arthropod vectors have multiple physical and immunological barriers that impede the development and transmission of parasites to new vertebrate hosts. These barriers include the peritrophic matrix (PM), a chitinous barrier that separates the blood bolus from the midgut epithelia and inturn, modulates vector-microbiota interactions. In tsetse flies, a sleeve-like PM is continuously produced by the cardia organ located at the fore- and midgut junction. African trypanosomes, Trypanosoma brucei, must bypass the PM twice; first to colonize the midgut and secondly to reach the salivary glands (SG), to complete their transmission cycle in tsetse. However, not all flies with midgut infections develop mammalian transmissible SG infections - the reasons for which are unclear. Here, we used transcriptomics, microscopy and functional genomics analyses to understand the factors that regulate parasite migration from midgut to SG. In flies with midgut infections only, parasites fail to cross the PM as they are eliminated from the cardia by reactive oxygen intermediates (ROIs) - albeit at the expense of collateral cytotoxic damage to the cardia. In flies with midgut and SG infections, expression of genes encoding components of the PM is reduced in the cardia, and structural integrity of the PM barrier is compromised. Under these circumstances trypanosomes traverse through the newly secreted and compromised PM. The process of PM attrition that enables the parasites to re-enter into the midgut lumen is apparently mediated by components of the parasites residing in the cardia. Thus, a fine-tuned dialogue between tsetse and trypanosomes at the cardia determines the outcome of PM integrity and trypanosome transmission success.\n\nAuthor summaryInsects are responsible for transmission of parasites that cause deadly diseases in humans and animals. Understanding the key factors that enhance or interfere with parasite transmission processes can result in new control strategies. Here, we report that a proportion of tsetse flies with African trypanosome infections in their midgut can prevent parasites from migrating to the salivary glands, albeit at the expense of collateral damage. In a subset of flies with gut infections, the parasites manipulate the integrity of a midgut barrier, called the peritrophic matrix, and reach the salivary glands for transmission to the next mammal. Either targeting parasite manipulative processes or enhancing peritrophic matrix integrity could reduce parasite transmission.

microbiology

Tandem repeats contribute to coding sequence variation in bumblebees (Hymenoptera: Apidae)

Tandem repeats (TRs) are highly dynamic regions of the genome. Mutations at these loci represent a significant source of genetic variation and can facilitate rapid adaptation. Bumblebees are important pollinating insects occupying a wide range of habitats. However, to date, molecular mechanisms underlying the potential adaptation of bumblebees to diverse habitats are largely unknown. In the present study, we investigate how TRs contribute to genetic variation in bumblebees, thus potentially facilitating adaptation. We identified 26,595 TRs in the buff-tailed bumblebee (Bombus terrestris) genome, 66.7% of which reside in genic regions. We also compared TRs found in B. terrestris with those present in the whole genome sequence of a congener, B. impatiens. We found that a total of 1,137 TRs were variable in length between the two sequenced bumblebee species, and further analysis reveals that 101 of them are located within coding regions. The 101 TRs were responsible for coding sequence variation and corresponded to protein sequence length variation between the two bumblebee species. The variability of identified TRs in coding regions between bumblebees was confirmed by PCR amplification of a subset of loci. Functional classification of bumblebee genes where coding sequences include variable-length TRs suggests that a majority of these genes are related to transcriptional regulation. Our results show that TRs contribute to coding sequence variation in bumblebees and TRs may facilitate the adaptation of bumblebees through diversifying proteins involved in controlling gene expression.

genomics

Multi-platform discovery of haplotype-resolved structural variation in human genomes

The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, and strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three human parent-child trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs ([&ge;]50 bp) per human genome. We also discover 156 inversions per genome--most of which previously escaped detection. Fifty-eight of the inversions we discovered intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The method and the dataset serve as a gold standard for the scientific community and we make specific recommendations for maximizing structural variation sensitivity for future large-scale genome sequencing studies.

genomics

A high-throughput analysis method of microdroplet PCR coupled with fluorescence spectrophotometry

Here we report a novel microdroplet PCR method combined with fluorescence spectrophotometry (MPFS), which allows for qualitative, quantitative and high -throughput detection of multiple DNA targets. In this study, each pair of primers was labeled with a specific fluorophore. Through microdroplet PCR, a target DNA was amplified and labeled with the same fluorophore. After products purification, the DNA products tagged with different fluorophores could be analyzed qualitatively by the fluorescent intensity determination. The relative fluorensence unit was also measured to construct the standard curve and to achieve quantitative analysis. In a reaction, the co -amplified products with different fluorophores could be simultaneously analyzed to achieve high -throughput detection. We used four kinds of GM maize as a model to confirm this theory. The qualitative results revealed high specificity and sensitivity of 0.5% (w / w). The quantitative results revealed that the limit of detection was 103copies and with good repeatability. Moreover, reproducibility assay were further performed using four foodborne pathogenic bacteria. Consequently, the same qualitative, quantitative and high-throughput results were confirmed as the four GM maize.

biochemistry

A plant receptor-like kinase promotes cell-to-cell spread of RNAi and is targeted by a virus

RNA interference (RNAi) in plants can move from cell to cell, allowing for systemic spread of an anti-viral immune response. How this cell-to-cell spread of silencing is regulated is currently unknown. Here, we describe that the C4 protein from Tomato yellow leaf curl virus has the ability to inhibit the intercellular spread of RNAi. Using this viral protein as a probe, we have identified the receptor-like kinase (RLK) BARELY ANY MERISTEM 1 (BAM1) as a positive regulator of the cell-to-cell movement of RNAi, and determined that BAM1 and its closest homologue, BAM2, play a redundant role in this process. C4 interacts with the intracellular domain of BAM1 and BAM2 at the plasma membrane and plasmodesmata, the cytoplasmic connections between plant cells, interfering with the function of these RLKs in the cell-to-cell spread of RNAi. Our results identify BAM1 as an element required for the cell-to-cell spread of RNAi and highlight that signalling components have been co-opted to play multiple functions in plants.

plant biology

A quantitative chemotherapy genetic interaction map reveals new factors associated with PARP inhibitor resistance

Nearly every cancer patient is treated with chemotherapy yet our understanding of factors that dictate response and resistance to such agents remains limited. We report the generation of a quantitative chemical-genetic interaction map in human mammary epithelial cells that charts the impact of knockdown of 625 cancer and DNA repair related genes on sensitivity to 29 drugs, covering all classes of cancer chemotherapeutics. This quantitative map is predictive of interactions maintained in cancer cell lines and can be used to identify new cancer-associated DNA repair factors, predict cancer cell line responses to therapy and prioritize drug combinations. We identify that GPBP1 loss in breast and ovarian cancer confers resistance to cisplatin and PARP inhibitors through the regulation of genes involved in homologous recombination. This map may help navigate patient genomic data and optimize chemotherapeutic regimens by delineating factors involved in the response to specific types of DNA damage.

cancer biology

Nucleosomes and DNA methylation shape meiotic DSB frequency in Arabidopsis transposons and gene regulatory regions

Meiotic recombination initiates via DNA double strand breaks (DSBs) generated by SPO11 topoisomerase-like complexes. Recombination frequency varies extensively along eukaryotic chromosomes, with hotspots controlled by chromatin and DNA sequence. To map meiotic DSBs throughout a plant genome, we purified and sequenced Arabidopsis SPO11-1-oligonucleotides. DSB hotspots occurred in gene promoters, terminators and introns, driven by AT-sequence richness, which excludes nucleosomes and allows SPO11-1 access. A strong positive relationship was observed between SPO11-1 DSBs and final crossover levels. Euchromatic marks promote recombination in fungi and mammals, and consistently we observe H3K4me3 enrichment in proximity to DSB hotspots at gene 5-ends. Repetitive transposons are thought to be recombination-silenced during meiosis, in order to prevent non-allelic interactions and genome instability. Unexpectedly, we found strong DSB hotspots in nucleosome-depleted Helitron/Pogo/Tc1/Mariner DNA transposons, whereas retrotransposons were coldspots. Hotspot transposons are enriched within gene regulatory regions and in proximity to immunity genes, suggesting a role as recombination-enhancers. As transposon mobility in plant genomes is restricted by DNA methylation, we used the met1 DNA methyltransferase mutant to investigate the role of heterochromatin on the DSB landscape. Epigenetic activation of transposon meiotic DSBs occurred in met1 mutants, coincident with reduced nucleosome occupancy, gain of transcription and H3K4me3. Increased met1 SPO11-1 DSBs occurred most strongly within centromeres and Gypsy and CACTA/EnSpm coldspot transposons. Together, our work reveals complex interactions between chromatin and meiotic DSBs within genes and transposons, with significance for the diversity and evolution of plant genomes.

genomics

Epigenetic activation of meiotic recombination in Arabidopsis centromeres via loss of H3K9me2 and non-CG DNA methylation

Eukaryotic centromeres contain the kinetochore, which connects chromosomes to the spindle allowing segregation. During meiosis centromeres are suppressed for crossovers, as recombination in these regions can cause chromosome mis-segregation. Plant centromeres are surrounded by repetitive, transposon-dense heterochromatin that is epigenetically silenced by histone 3 lysine 9 dimethylation (H3K9me2), and DNA methylation in CG and non-CG sequence contexts. Here we show that disruption of Arabidopsis H3K9me2 and non-CG DNA methylation pathways increases meiotic DNA double strand breaks (DSBs) within centromeres, whereas crossovers increase within pericentromeric heterochromatin. Increased pericentromeric crossovers in H3K9me2/non-CG mutants occurs in both inbred and hybrid backgrounds, and involves the interfering crossover repair pathway. Epigenetic activation of recombination may also account for the curious tendency of maize transposon Ds to disrupt CHROMOMETHYLASE3 when launched from proximal loci. Thus H3K9me2 and non-CG DNA methylation exert differential control of meiotic DSB and crossover formation in centromeric and pericentromeric heterochromatin.

genomics

Isolation And Characterization Of Key Genes That Promote Flavonoid Accumulation In Purple-Leaf Tea (Camellia sinensis L.)

There were several high concentrations of flavonoid components in tea leaves that present health benefits. A novel purple-leaf tea variety, Mooma1, was obtained from the natural hybrid population of Longjing 43 variety. The buds and young leaves of Mooma1 were displayed in bright red. HPLC and LC-MS analysis showed that anthocyanins and O-Glycosylated flavonols were remarkably accumulated in the leaves of Mooma1, while the total amount of catechins in purple-leaf leaves was slightly decreased compared with the control. A R2R3-MYB transcription factor (CsMYB6A) and a novel UGT gene (CsUGT72AM1), that were highly expressed in purple leaf were isolated and identified by transcriptome sequencing. The over-expression of transgenic tobacco confirmed that CsMYB6A can activate the expression of flavonoid-related structural genes, especially CHS and 3GT, controlling the accumulation of anthocyanins in the leaf of transgenic tobacco. Enzymatic assays in vitro confirmed that CsUGT72AM1 has catalytic activity as a flavonol 3-O-glucosyltransferase, and displayed broad substrate specificity. The results were useful for further elucidating the molecular mechanisms of the flavonoid metabolic fluxes in the tea plant.

plant biology

VaPoR: a high-speed validation approach for structural variation using long-read sequencing technology.

AbstractSummaryAlthough there are numerous algorithms that have been developed to identify structural variation (SVs) in genomic sequences, there is a dearth of approaches that can be used to evaluate their results. The emergence of new sequencing technologies that generate longer sequence reads can, in theory, provide direct evidence for all types of SVs regardless of the length of region through which it spans. However, current efforts to use these data in this manner require the use of large computational resources to assemble these sequences as well as manual inspection of each region. Here, we present VaPoR, a highly efficient algorithm that autonomously validates large SV sets using long read sequencing data. We assess of the performance of VaPoR on both simulated and real SVs with regards to various features including accuracy and sensitivity of breakpoint evaluation and report a high fidelity rate.\n\nAvailabilityhttps://github.com/mills-lab/VaPoR\n\nContactremills@umich.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

A Network Integration Approach for Drug-Target Interaction Prediction and Computational Drug Repositioning from Heterogeneous Information

The emergence of large-scale genomic, chemical and pharmacological data provides new opportunities for drug discovery and repositioning. Systematic integration of these heterogeneous data not only serves as a promising tool for identifying new drug-target interactions (DTIs), which is an important step in drug development, but also provides a more complete understanding of the molecular mechanisms of drug action. In this work, we integrate diverse drug-related information, including drugs, proteins, diseases and side-effects, together with their interactions, associations or similarities, to construct a heterogeneous network with 12,015 nodes and 1,895,445 edges. We then develop a new computational pipeline, called DTINet, to predict novel drug-target interactions from the constructed heterogeneous network. Specifically, DTINet focuses on learning a low-dimensional vector representation of features for each node, which accurately explains the topological properties of individual nodes in the heterogeneous network, and then predicts the likelihood of a new DTI based on these representations via a vector space projection scheme. DTINet achieves substantial performance improvement over other state-of-the-art methods for DTI prediction. Moreover, we have experimentally validated the novel interactions between three drugs and the cyclooxygenase (COX) protein family predicted by DTINet, and demonstrated the new potential applications of these identified COX inhibitors in preventing inflammatory diseases. These results indicate that DTINet can provide a practically useful tool for integrating heterogeneous information to predict new drug-target interactions and repurpose existing drugs. The source code of DTINet and the input heterogeneous network data can be downloaded from http://github.com/luoyunan/DTINet.

bioinformatics