Search bioRxiv⌕ Search

Biology subjects

Dou, F.

Publications and source records attributed to Dou, F..

2 recordsLinked to original sources

Glydentify: An explainable deep learning platform for glycosyltransferase donor substrate prediction

Glycosyltransferases (GTs) are a large family of enzymes that catalyze glycosidic linkages formation between chemically diverse donor and acceptor molecules to regulate diverse cellular processes across all domains of life. Despite their importance, the activated sugar donors (donor substrates) used by most GTs remain unidentified, limiting our understanding of GT functions. To address this challenge, we developed Glydentify, a deep learning framework that predicts donor usage across GT-A and GT-B fold glycosyltransferases. Trained on large-scale UniProt annotations, Glydentify integrates protein sequence embeddings learned from protein language models with chemical features derived from molecular encoders trained on extensive chemical datasets. The resulting models achieve high predictive performance, with precision-recall AUCs (PR-AUC) of 0.86 for GT-A and 0.91 for GT-B, surpassing general enzyme-substrate predictors while requiring minimal manual curation. We employed Glydentify to predict the donor specificity of uncharacterized plant GTs and experimentally tested the predictions using in vitro biochemical assays. Furthermore, we demonstrate that the model utilizes a combination of evolutionary, structural, and biochemical features to predict donor specificity through residue attention score analysis. Together, these results establish Glydentify as a robust, explainable framework for decoding donor-glycosyltransferase relationships and highlight its potential as a broadly applicable framework for modeling enzyme classes that act on chemically diverse substrates.

bioinformatics↗

Identification of Differentially Expressed Genes and Proteins Related to Diapause in Lymantria Dispar: Insights for the Mechanism of Diapause from Transcriptome and Proteome Analyses

Spongy moth (Lymantria dispar Linnaeus) is a globally recognized quarantine leaf-eating pest. Spongy moths typically enter diapause after completing embryonic development and overwinter in the egg stage. They spend three-quarters of their life cycle (approximately nine months) in the egg stage, which requires a period of low-temperature stimulation to break diapause and continue growth and development. In this study, we explored the molecular mechanism underlying the diapause process in spongy moth. We performed bioinformatics analysis on four Asian populations of spongy moth and one Asian-European hybrid population through a transcriptome analysis combined with proteomics. The results revealed that 1,842 genes were differentially expressed upon diapause initiation, while 264 genes were identified upon diapause termination. Eight diapause-related genes were screened out from the three-level pathways that were significantly enriched by differentially expressed genes at the time of diapause and diapause termination, and the phylogenetic tree and protein three-dimensional structure model were constructed. This study elucidates the diapause mechanism of spongy moth at the gene and protein levels, providing theoretical insights into the early and precise prevention and control of spongy moth. This study can facilitate the development of an efficient, environmentally friendly control system for managing spongy moth populations in the field.

genomics↗