Search bioRxiv⌕ Search

Biology subjects

Bocker, S.

Publications and source records attributed to Bocker, S..

2 recordsLinked to original sources

MAD HATTER correctly annotates 98% of small molecule tandem mass spectra searching in PubChem

Metabolites provide a direct functional signature of cellular state. Untargeted metabolomics usually relies on mass spectrometry, a technology capable of detecting thousands of compounds in a biological sample. Metabolite annotation is executed using tandem mass spectrometry. Spectral library search is far from comprehensive, and numerous compounds remain unannotated. So-called in silico methods allow us to overcome the restrictions of spectral libraries, by searching in much larger molecular structure databases. Yet, after more than a decade of method development, in silico methods still do not reach correct annotation rates that users would wish for. Here, we present a novel computational method called MO_SCPLOWADC_SCPLOW HO_SCPLOWATTERC_SCPLOW for this task. MO_SCPLOWADC_SCPLOW HO_SCPLOWATTERC_SCPLOW combines CSI:FingerID results with information from the searched structure database via a metascore. Compound information includes the melting point, and the number words in the compound description starting with the letter u. We then show that MO_SCPLOWADC_SCPLOW HO_SCPLOWATTERC_SCPLOW reaches a stunning 97.6% correct annotations when searching PubChem, one of the largest and most comprehensive molecular structure databases. Finally, we explain what evaluation glitches were necessary for MO_SCPLOWADC_SCPLOW HO_SCPLOWATTERC_SCPLOW to reach this annotation level, what is wrong with similar metascores in general, and why metascores may screw up not only method evaluations but also the analysis of biological experiments.

bioinformatics↗

Chemically-informed Analyses of Metabolomics Mass Spectrometry Data with Qemistree

Untargeted mass spectrometry is employed to detect small molecules in complex biospecimens, generating data that are difficult to interpret. We developed Qemistree, a data exploration strategy based on hierarchical organization of molecular fingerprints predicted from fragmentation spectra, represented in the context of sample metadata and chemical ontologies. By expressing molecular relationships as a tree, we can apply ecological tools, designed around the relatedness of DNA sequences, to study chemical composition.

bioinformatics↗