Search bioRxivSearch

EXPLORE THE ARCHIVE

Wang, J.

Publications and source records attributed to Wang, J..

4 recordsLinked to original sources

High-Resolution Subtyping of Pediatric Low-Grade Glioma Using an Integrated Meta-Clustering Framework

Pediatric low-grade glioma (pLGG) is the most common type of brain tumor in children, accounting for approximately 30% of all central nervous system tumors in children. pLGG has multiple molecular subtypes that differ in disease progression, recurrence patterns, and treatment responses. Conventional wet lab approaches including molecular profiling and histopathological studies for pLGG characterization are time consuming, costly, and laborious. Recently, methods based on artificial intelligence (AI) or machine learning (ML) have been widely used for pLGG molecular categorization, but most of them can only identify two or three pLGG subtypes. To more comprehensively characterize the molecular subtypes of pLGG and their potential biological and therapeutic significance, we develop an integrated meta-clustering approach, namely Meta-pLGG, that can explore high resolution molecular subtypes and their transcriptional heterogeneity for pLGG. Specifically, we first performed multiple rounds of random projection (RP) to generate dimension-reduced feature vectors from pLGG transcriptomics data, each of which was subsequently clustered by different clustering algorithms including hierarchical clustering, K-means, Self-Organizing Maps (SOM), Non-negative Matrix Factorization (NMF), Gaussian Mixture Model (GMM), and Spectral Clustering, as base clustering methods. Then, to yield robust clustering performance, we integrated the clustering results of these RP based individual clustering algorithms by adopting a weighted meta-clustering (wMetaC) approach. Results based on 532 pLGG patients suggested that our proposed approach demonstrated superior stability and discriminative powers for higher resolution pLGG subtyping compared to conventional approaches. Based on consensus matrix analysis, we identified two major pLGG mega-subtypes, with one further subdivided into three subgroups and the other into two. Then, we performed cluster specific differential gene expression analysis, molecular pathway analysis, and gene-drug-disease association analysis. The results showed that the identified five subgroups exhibited significant subtype-specific transcriptomic heterogeneity. In summary, our meta-clustering approach demonstrated much higher performance and robustness in identifying higher resolution molecular subtypes of pLGG, revealing the molecular heterogeneity within pLGG and potentially providing new insights for more precise molecular subtyping and precision therapy.

bioinformatics

Detection of Frustration-related Operant Behavior in Rats via Machine Learning Methods

Despite its strong link to neuropsychiatric conditions, frustration remains critically understudied in humans and animals alike. Therefore, there is an urgent need to develop tools to understand and therapeutically target frustration-related functions. Interestingly, humans and rats respond similarly during frustrative nonreward by increasing barpress durations. We previously validated barpress duration in rat operant tasks as a reliable measure of frustration-related behavior; however, it is wellknown that in addition to duration of responding, emotional states such as frustration alter other aspects of responding such as force of pressing. One-dimensional, static measures such as maximum force could miss rich information contained within operant data. Thus, the objective of this study is to apply machine learning (ML) to force/time profiles to discriminate frustration-related barpresses from non-frustration-related barpresses. Results showed an AUROC for FR1 (i.e., non-frustrated) vs. extinction (frustrated condition) for individual barpresses of 0.65 that improved to 0.84 with a chunk size of 10. The model generalized well to progressive ratio responding, a different kind of frustration procedure. We conclude that force/time profiling does provide utility beyond one dimensional measures of duration or force separately, meaning that we can indeed infer the internal state of frustration from behavior using ML techniques. Importantly, this project will also serve as proof-of-concept for applying ML to predict other internal states from barpress data.

animal behavior and cognition

Dissecting the TMEM132A-EGFR Dependency to Unlock Translational Therapeutic Opportunities for Pan-Solid Tumor

Solid tumors remain refractory to conventional treatments, yet cell surface proteins, by virtue of their extracellular accessibility and critical roles in tumor signaling, represent an attractive class of targets for precision-targeted therapy. Here, we report that TMEM132A is an essential and previously unrecognized pan-cancer target. TMEM132A interacts directly with EGFR and stabilizes its expression, thereby tethering EGFR at the plasma membrane and sustaining constitutive activation of lipid synthesis. Mechanistically, the TMEM132A-EGFR axis promotes lipogenesis by facilitating SREBP nuclear translocation, which in turn upregulates ACLY and ACSS2 expression to drive acetyl-CoA production and downstream lipid biosynthesis, ultimately disrupting lipid droplet homeostasis. To therapeutically target this axis, we developed a nanobody, LFNanoT132A#3, which effectively blocks the TMEM132A-EGFR interaction, abrogates downstream signaling activation, and potently inhibits proliferation across multiple solid tumor types. Notably, LFNanoT132A also exerts robust antitumor activity against H1975 xenografts, a model resistant to first- and second- generation EGFR inhibitors, underscoring its potential to overcome conventional drug resistance. Our findings establish TMEM132A#3 as a critical node in membrane-tethered oncogenic signaling and metabolic rewiring, and position LFNanoT132A#3 as a promising therapeutic candidate for precision cancer therapy.

cancer biology

TomatoPGFM: A graph-conditioned foundation model for tomato pangenomes

Most genomic foundation models are pretrained on independent linear assemblies and therefore do not explicitly represent population-level segment sharing or local graph connectivity. We developed TomatoPGFM, a graph-conditioned model pretrained on 54.65 Gb of sequence from 66 tomato (Solanum spp.) accessions. Sequence tokens were conditioned on pangenome node attributes and local adjacency, and the model was optimised using masked language modelling and graph-feature reconstruction. To evaluate model responses to graph-conditioned input, we compared aligned, shuffled and disabled graph inputs in 25,000 windows from the training panel. Sequence-aligned graph input produced lower masked language modelling loss than graph-off at all five curriculum stages in both training-panel strata, while the shuffled perturbation generally yielded intermediate losses. We then assessed sequence-only transfer in Solanum sitiens LA1974 and S. lycopersicum MicroTom, neither of which was used for graph construction or pretraining. Frozen-probe AUROC values for gene-versus-intergenic and coding-sequence-versus-intergenic classification ranged from 0.8489 to 0.9593. TomatoPGFM produced higher AUROC point estimates than DNABERT-2 in all four comparisons. Enabling the zero-feature GraphAdapter pathway with adjacency messaging disabled changed throughput by less than 1% at 512-2,048 positions under the tested configuration. Together, these results show that TomatoPGFM responds consistently to sequence-aligned pangenome context in training-panel sequences and provides informative sequence representations for genic-region classification in accessions excluded from graph construction and pretraining.

bioinformatics