Search bioRxiv⌕ Search

Biology subjects

Zhou Chen, Y.

Publications and source records attributed to Zhou Chen, Y..

2 recordsLinked to original sources

FLARE: Fine-grained Learning for Alignment of spectra-molecule REpresentation Enhances Metabolite Annotation

Accurate metabolite annotation via tandem mass spectrometry remains a major bottleneck in untargeted metabolomics. Recent implicit models that avoid molecular generation or spectra simulation have shown competitive performance by aligning spectra and molecular structures in the embedding space. Still, they overlook the detailed relationships between spectral peaks and molecular substructures that govern fragmentation. We introduce FLARE (Fine-grained Learning for Alignment of spectra-molecule REpresentations), a contrastive learning framework that leverages bidirectional peak-node alignment under learned weak supervision. Unlike models that rely solely on global embeddings, FLARE computes similarity via maxima over peak-to-atom and atom-to-peak interactions, capturing chemically meaningful local correspondences and enabling interpretable spectra-molecule matching. It achieves state-of-the-art results on MassSpecGym, with 43.15% rank@1 (mass-based) and 22.66% (formula-based), surpassing previous models by over 63%. FLAREs learned embeddings correspond with molecular classes, match fingerprint similarity, and detect differential metabolites in a breast cancer xenograft study, showcasing its translational potential.

systems biology↗

Learning from All Views: A Multiview Contrastive Framework for Metabolite Annotation

Metabolomics, enabled by high-throughput mass spectrometry, promises to advance our understanding of cellular biochemistry and guide new discoveries in disease mechanisms, drug development, and personalized medicine. However, as the assignment of molecular structures to measured spectra is challenging, annotation rates remain low and hinder potential advancements. We present MultiView Projection (MVP), a novel framework for learning a joint embedding space between molecules and spectra by leveraging multiple data views: molecular graphs, molecular fingerprints, spectra, and consensus spectra. MVP builds on contrastive multiview learning to capture mutual information across views, leading to more robust and generalizable representations for spectral annotation. Unlike prior approaches that consider multiple views via concatenation or as targets of auxiliary tasks, MVP learns from all views jointly, resulting in improved molecular candidate ranking. Notably, MVP supports annotation using either individual spectra or consensus spectra, enabling flexible use of multiple measurements. On the MassSpecGym benchmark, we show that annotation using query consensus spectra significantly outperforms rank aggregation strategies based on constituent spectrum annotation. Using the consensus spectrum view, MVP achieves 35.99% and 13.96% rank@1 when retrieving candidates by mass and formula, respectively. When ranking using individual spectra, MVP demonstrates performance that is superior to or on par with existing methods, achieving 26.37% and 11.10% rank@1 for candidates by mass and formula, respectively. MVP offers a flexible, extensible foundation for learning from multiple molecule/spectra data views. For Table of Contents Only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/688047v1_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@f6c219org.highwire.dtl.DTLVardef@40f8d4org.highwire.dtl.DTLVardef@19031b6org.highwire.dtl.DTLVardef@1afa58d_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗