Search bioRxiv⌕ Search

Biology subjects

Kaixuan, D.

Publications and source records attributed to Kaixuan, D..

2 recordsLinked to original sources

PlantGeneAnn: a strand-specific genome foundation model for ab initio gene structure annotation of plant genomes

High-quality plant genome assemblies are rapidly increasing, but accurate structural annotation remains reliant on transcript and homology evidence, limiting applications in newly sequenced and non-model species. Here, we present PlantGeneAnn, a plant-optimized, strand-specific genome foundation model for ab initio gene structure annotation. Fine-tuned on only nine high-quality model plant annotations, PlantGeneAnn outperformed a multi-species model trained on 42 species, showing that annotation quality is more important than token volume. On a stringent 13-species benchmark covering rosids, asterids, and monocots, PlantGeneAnn surpassed four state-of-the-art baselines across five evaluation levels, from base-level classification to complete transcript recovery. It achieved higher intron precision and better captured complex gene structures. In zero-shot variant effect prediction, PlantGeneAnn identified cryptic splice donors and premature stop codons in maize and rice, with saturation mutagenesis confirming single-nucleotide, context-dependent sensitivity. It also retained generalizability for epigenomic track prediction, highlighting its value for pan-genomics, crop improvement, and non-model plant research.

bioinformatics↗

Precise metabolic dependencies of cancer through deep learning and validations

Cancer cells exhibit metabolic reprogramming to sustain proliferation, creating metabolic vulnerabilities absent in normal cells. While prior studies identified specific metabolic dependencies, systematic insights remain limited. Here we build a graph deep learning based metabolic vulnerability prediction model "DeepMeta", which can accurately predict the dependent metabolic genes for cancer samples based on transcriptome and metabolic network information. The performance of DeepMeta has been extensively validated with independent datasets. The metabolic vulnerability of "undruggable" cancer driving alterations have been systematically explored using the cancer genome atlas (TCGA) dataset. Notably, CTNNB1 T41A activating mutations showed experimentally confirmed vulnerability to purine/pyrimidine metabolism inhibition. TCGA patients with the predicted pyrimidine metabolism dependency show a dramatically improved clinical response to chemotherapeutic drugs that block this pyrimidine metabolism pathway. This study systematically uncovers the metabolic dependency of cancer cells, and provides metabolic targets for cancers driven by genetic alterations that are originally undruggable on their own.

cancer biology↗