Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.07.30.741704

Quantifying per-match Reliability in Library Matching for Untargeted Metabolomics Workflows

Abstract

Tandem mass spectrometry has become central to untargeted metabolomics. The translation of unknown spectra into biological insight depends on assigning chemical identities to detected metabolites. Structural characterization typically begins with mass spectral library matching, in which experimental spectra are compared against reference libraries and candidate annotations are ranked by their spectral similarity to the query. As spectral libraries and experimental datasets grow, however, more candidates achieve comparable similarity scores for a single query, and similarity scores give no indication of how reproducible a candidate match is or how sensitive it is to the underlying fragment evidence. Existing false-discovery-rate approaches can indicate annotation error at the dataset level but do not provide a per-match estimate of reliability. Here, we introduce a SpecReBoot-inspired query-focused bootstrapping approach that resamples the fragment evidence of each query spectrum. This approach relies on recomputing query similarity to candidate library spectra across bootstrap replicates, which provides a statistical distribution of scores rather than a single value. From this distribution we define the match support, a per-match reliability estimate quantifying the reproducibility of a match under spectral perturbation, together with measures of ranking stability that describe how often a candidate remains among the top-ranked matches across replicates. Applied to a forensic drug-of-abuse case, match support distinguished previously identified annotations from high-scoring false positives: a distinction cosine similarity failed to make. Furthermore, match support values remained stable as the reference library was expanded, whereas ranking stability metrics shifted significantly. In a cross-instrument endogenous metabolite library search, match support further revealed metric-specific annotation behavior, identifying metabolites consistently supported across different similarity metrics, while flagging annotations whose reliability depended strongly on the chosen scoring metric. Benchmarking against a natural-product reference library demonstrated that ranking based on match support values promoted true matches by four ranks on average compared with cosine-based ranking, without promoting analogs. Under controlled spectral perturbation experiments, match support flagged incorrect annotations with an AUROC of 0.75, whereas the cosine similarity score alone of the same match reached only 0.56. Query-focused bootstrapping thus provides a practical, per-match measure of annotation reliability, bringing the field a step toward reliable annotations at scale. We anticipate that incorporation of our annotation reliability scoring into computational metabolomics workflows will further promote the growth of spectral libraries and enhance their applicability across scientific disciplines.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Charria-Giron, E., van IJcken, J., Della Vedova, L., Torres-Ortega, L. R., van der Hooft, J. J. J.. 2026-07-30. Quantifying per-match Reliability in Library Matching for Untargeted Metabolomics Workflows. https://doi.org/10.64898/2026.07.30.741704

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Deep reinforcement learning-driven discovery of a MsbA-targeted small-molecule antibiotic for the treatment of Acinetobacter baumannii infection

Antibiotics with new mechanisms are highly pursued to address the threat of infections caused by drug-resistant Gram-negative bacteria. Targeting MsbA, a key protein of the lipopolysaccharide biosynthesis pathway, represents a promising strategy to discover new classes of antibiotics. However, currently available MsbA-targeted molecules either lack sufficient potency or have unfavorable properties, necessitating expansion of chemical space. In this study, we chose the most promising cerastecin Cpd 4 as the template, and used two Artificial Intelligence (AI)-based tools, i.e. Link-INVENT and AutoMolDesigner for molecular design, performed chemical derivatization and antibacterial activity evaluation, which led to the discovery of Y-11 (MIC for A. baumannii: 0.5 g/mL). Encouragingly, Y-11 showed equivalent potency to Cpd4 for carbapenem-resistant A. baumannii, and less cytotoxicity and hemolysis as well as lower spontaneous resistance frequency. In vivo efficacy study demonstrated that Y-11 could effectively reduce bacterial loads in the mice infected by A. baumannii. The following mechanism study including molecular dynamics simulation, biochemical assay, and transmission electron microscope (TEM) analysis suggested that Y-11 inhibited the transport of lipooligosaccharide and impaired the formation of outer membrane, probably by competitively binding to the substrate binding site of MsbA and modulating ATPase activity. Taken together, we have discovered a MsbA-targeted small molecule Y-11 via AI-driven drug design, which provides a foundation for future antibiotic development.

biochemistry↗

Dynamic architecture of the Rixosome reveals mechanism of activation and ITS2 processing

Eukaryotic ribosome assembly requires the coordinated processing and extensive remodeling of pre-rRNAs. During late nuclear maturation of the 60S subunit, sequential removal of the internal transcribed spacer 2 (ITS2) is initiated by endonucleolytic cleavage at site C2 by the conserved Las1 nuclease. Las1 acts together with the kinase Grc3 and the Rix1 complex to form the Rixosome, which also functions in transcriptional regulation. However, the assembly of the Rixosome, its recruitment to pre-ribosomes, and its activation for ITS2 cleavage remain unclear. Here, we present cryo-EM structures of the human LAS1 complex, two structures of the isolated Rixosome and nine transition states of Rix1-bound pre-60S particles from Schizosaccharomyces pombe. These structures reveal a dynamic Rixosome architecture in which the heterotetrameric Las1 complex engages one or two copies of the Rix1 complex. Rix1 binding is highly flexible in the human Rixosome but rigid in the yeast complex. The isolated yeast Rixosome remains inactive, but binding to the pre-60S particle triggers a structural rearrangement that allows for substrate engagement and activation of the nuclease. Together, our results define the dynamic architecture of the Rixosome and provide a structural framework for ITS2 processing during nuclear maturation of the eukaryotic 60S ribosomal subunit.

biochemistry↗

SGFP-Grid Split GFP Graphene Grids

Affinity graphene grids provide a promising approach for selective protein capture in cryo-EM. Here, we introduce a split-GFP graphene grid platform(SGFP-G), in which graphene-conjugated GFP 1-10 selectively captures GFP11 tagged proteins from low concentration samples or cell lysates. This platform enables rapid assessment of target protein enrichment and particle distribution before vitrification via fluorescence imaging, while the grid design positions captured proteins away from the graphene surface and air-water interface. We also introduce a unique strategy to minimize nonspecific protein adsorption, thereby improving the selective enrichment of target proteins on this grid. Using GFP11-tagged apoferritin, we demonstrate fluorescence guided protein capture and obtain a 2.58 [A] cryoEM reconstruction, establishing SGFP-G as an affinity grid platform for high resolution structural studies with reduced sample requirements.

biochemistry↗