Search bioRxivSearch

Biology subjects

Ye, W.

Publications and source records attributed to Ye, W..

7 recordsLinked to original sources

Automatic Human-like Mining and Constructing Reliable Genetic Association Database with Deep Reinforcement Learning

The increasing amount of scientific literature in biological and biomedical science research has created a challenge in the continuous and reliable curation of the latest knowledge discovered, and automatic biomedical text-mining has been one of the answers to this chal-lenge. In this paper, we aim to further improve the reliability of biomedical text-mining by training the system to directly simulate the human behaviors such as querying the PubMed, selecting articles from queried results, and reading selected articles for knowledge. We take advantage of the efficiency of biomedical text-mining, the flexibility of deep reinforcement learning, and the massive amount of knowledge collected in UMLS into an integrative arti-ficial intelligent reader that can automatically identify the authentic articles and effectively acquire the knowledge conveyed in the articles. We construct a system, whose current pri-mary task is to build the genetic association database between genes and complex traits of the human. Our contributions in this paper are three-fold: 1) We propose to improve the reliability of text-mining by building a system that can directly simulate the behavior of a researcher, and we develop corresponding methods, such as Bi-directional LSTM for text mining and Deep Q-Network for organizing behaviors. 2) We demonstrate the effec-tiveness of our system with an example in constructing a genetic association database. 3) We release our implementation as a generic framework for researchers in the community to conveniently construct other databases.

bioinformatics

Genome-wide identification and expression specificity analysis of the DNA methyltransferase gene family under adversity stresses in cotton

DNA methylation is an important epigenetic mode of genomic DNA modification that is an important part of maintaining epigenetic content and regulating gene expression. DNA methyltransferases (MTases) are the key enzymes in the process of DNA methylation. Thus far, there has been no systematic analysis the DNA MTases found in cotton. In this study, the whole genome of cotton C5-Mtase coding genes was identified and analyzed using a bioinformatics method based on information from the cotton genome. In this study, 51 DNA MTase genes were identified, of which 8 belonged to G. raimondii (group D), 9 belonged to G. arboretum L. (group A), 16 belonged to G. hirsutum L. (group AD1) and 18 belonged to G. barbadebse L. (group AD2). Systematic evolutionary analysis divided the 51 genes into four subfamilies, including 7 MET homologous proteins, 25 CMT homologous proteins, 14 DRM homologous proteins and 5 DNMT2 homologous proteins. Further studies showed that the DNA MTases in cotton were more phylogenetically conserved. The comparison of their protein domains showed that the C-terminal functional domain of the 51 proteins had six conserved motifs involved in methylation modification, indicating that the protein has a basic catalytic methylation function and the difference in the N-terminal regulatory domains of the 51 proteins divided the proteins into four classes, MET, CMT, DRM and DNMT2, in which DNMT2 lacks an N-terminal regulatory domain. Gene expression in cotton is not the same under different stress treatments. Different expression patterns of DNA MTases show the functional diversity of the cotton DNA methyltransferase gene family. VIGS silenced Gossypium hirsutum l. in the cotton seedling of DNMT2 family gene GhDMT6, after stress treatment the growth condition was better than the control. The distribution of DNA MTases varies among cotton species. Different DNA MTase family members have different genetic structures, and the expression level changes with different stresses, showing tissue specificity. Under salt and drought stress, G. hirsutum L. TM-1 increased the number of genes more than G. raimondii and G. arboreum L. Shixiya 1. The resistance of Gossypium hirsutum L.TM-1 to cold, drought and salt stress was increased after the plants were silenced with GhDMT6 gene.

genomics

Localization and protein-protein interaction of protein kinase CK2 suggest a chaperone-like activity is integral to its function in M. oryzae.

Magnaporthe oryzae (Mo) is a model pathogen causing rice blast resulting in yield and economic losses world-wide. CK2 is a constitutively active, serine/threonine kinase in eukaryotes, having a wide array of known substrates and involved in many cellular processes. We investigated the localization and role of MoCK2 during growth and infection. BLAST search for MoCK2 components and targeted deletion of subunits was combined with protein-GFP fusions to investigate localization. We found one CKa and two CKb subunits of the CK2 holoenzyme. Deletion of the catalytic subunit CKa was not possible and might indicate that such deletions are lethal. The CKb subunits could be deleted but they were both necessary for normal growth and pathogenicity. Localization studies showed that the CK2 holoenzyme needed to be intact for normal localization at septal pores and at appressorium penetration pores. Nuclear localization of CKa was however not dependent on the intact CK2 holoenzyme. In appressoria, CK2 formed a large ring perpendicular to the penetration pore and the ring formation was dependent on the presence of all CK2 subunits. The effects on growth and pathogenicity of deletion of the b subunits combined with the localization indicate that CK2 can have important regulatory functions not only in the nucleus/nucleolus but also at fungal specific structures as septa and appressorial pores.

cell biology

Tumour purity as a prognostic factor in colon cancer

Tumour purity is defined as the proportion of cancer cells in the tumour tissue. The impact of tumour purity on colon cancer (CC) prognosis, genetic profile and microenvironment has not been thoroughly accessed. Therefore, clinical and transcriptomic data from three public datasets, GSE17536/17537, GSE39582, and TCGA were retrospectively collected (n = 1248). Tumour purity of each sample was inferred by a computational method based on transcriptomic data. Stage III and MMR-deficient (dMMR) CC patients showed a significantly lower tumour purity. Low purity CC conferred worse survival and tumour purity was identified as an independent prognostic factor. Moreover, high tumour purity CC patients benefited more from adjuvant chemotherapy. Subsequent genomic analysis found that the mutation burden was negatively associated with tumour purity with only APC and KRAS significantly more mutated in high purity CC. However, no somatic copy number alteration event was correlated with tumour purity. Furthermore, immune-related pathways and immunotherapy-associated markers (PD-1, PD-L1, CTLA-4, LAG-3, and TIM-3) were highly enriched in low purity samples. Notably, the relative proportion of M2 macrophages and neutrophils, which indicated worse survival in CC, was negatively associated with tumour purity. Therefore, tumour purity exhibited potential value for CC prognostic stratification as well as adjuvant chemotherapy benefit prediction. The relative worse survival in low purity CC may attribute to higher mutation frequency in key pathways and purity related microenvironmental changing.\n\nSummaryLow purity colon cancer patients conferred worse survival and benefited less from adjuvant chemotherapy. The mutation burden was negatively associated with tumour purity. Low purity samples exhibited intense immune phenotype with more M2 macrophages and neutrophils infiltration.

cancer biology

Phytophthora methylomes modulated by expanded 6mA methyltransferases are associated with adaptive genome regions

Filamentous plant pathogen genomes often display a bipartite architecture with gene sparse, repeat-rich compartments serving as a cradle for adaptive evolution. However, the extent to which this \"two-speed\" genome architecture is associated with genome-wide epigenetic modifications is unknown. Here, we show that the oomycete plant pathogens Phytophthora infestans and Phytophthora sojae possess functional adenine N6- methylation (6mA) methyltransferases that modulate patterns of 6mA marks across the genome. In contrast, 5-methylcytosine (5mC) could not be detected in the two Phytophthora species. Methylated DNA IP Sequencing (MeDIP-seq) of each species revealed that 6mA is depleted around the transcriptional starting sites (TSS) and is associated with low expressed genes, particularly transposable elements. Remarkably, genes occupying the gene-sparse regions have higher levels of 6mA compared to the remainder of both genomes, possibly implicating the methylome in adaptive evolution of Phytophthora. Among three putative adenine methyltransferases, DAMT1 and DAMT3 displayed robust enzymatic activities. Surprisingly, single knockouts of each of the 6mA methyltransferases in P. sojae significantly reduced in vivo 6mA levels, indicating that the three enzymes are not fully redundant. MeDIP-seq of the damt3 mutant revealed uneven patterns of 6mA methylation across genes, suggesting that PsDAMT3 may have a preference for gene body methylation after the TSS. Our findings provide evidence that 6mA modification is an epigenetic mark of Phytophthora genomes and that complex patterns of 6mA methylation by the expanded 6mA methyltransferases may be associated with adaptive evolution in these important plant pathogens.

molecular biology

Incidental identification of maternal malignancies in two Asian women underwent noninvasive prenatal test

Noninvasive prenatal test (NIPT) has been widely used as a screening test for trisomy 13, 18 and 21 worldwide. Recently, coexistence of maternal malignancy and pregnancy has drawn increasing attention in NIPT studies. Malignancy in pregnant women potentially affected NIPT results, which may cause false positive results or failed tests. However, no such case has ever been reported in Asian population. In this study, for the first time, we reported a stage III dysgerminoma and advanced gastric cancer during pregnancy accidentally identified during NIPT tests. These two women showed aberrant chromosome aneuploidies in NIPT results and concordant pattern of genome disruption found in tumor samples. The findings in this study further validate the effect of maternal malignancy on NIPT results and strengthen the possibility of detecting malignant tumors through NIPT in the future.

cancer biology

EuMicrobedbLite: A lightweight genomic resource and analytic platform for draft oomycete genomes

We have developed EuMicrobedbLite - A light weight comprehensive genome resource and sequence analysis platform for oomycete organisms. EuMicrobedbLite is a successor of the VBI Microbial Database (VMD) that was built using the Genome Unified Schema (GUS). In this version, the GUS schema has been greatly simplified with removal of many obsolete modules and redesign of others to incorporate contemporary data. Several dependencies such as perl object layers used for data loading in VMD have been replaced with independent light weight scripts. EumicrobedbLite now runs on a powerful annotation engine developed at our lab called \"Genome Annotator Lite\". Currently this database has 26 publicly available genomes and 10 EST datasets of oomycete organisms. The browser page has dynamic tracks presenting comparative genomics analyses, coding and non-coding data, tRNA genes, repeats and EST alignments. In addition, we have defined 44,777 core conserved proteins from twelve oomycete organisms that form 2974 clusters. Synteny viewing is enabled by incorporation of the Genome Synteny Viewer (GSV) tool. The user interface has undergone major changes for ease of browsing. Queryable comparative genomics information, conserved orthologous genes and pathways are among the new key features updated in this database. The browser has been upgraded to enable user upload of GFF files for quick view of genome annotation comparisons. The toolkit page integrates the EMBOSS package and has a gene prediction tool. Annotations for the organisms are updated once every six months to ensure quality. The database resource is available at www.eumicrobedb.org.

bioinformatics