bioRxiv · 10.1101/2020.02.06.937870
A Pre-computed Probabilistic Molecular Search Engine for Tandem Mass Spectrometry Proteomics
Abstract
Mass spectrometry methods of peptide identification involve comparing observed tandem spectra with in-silico derived spectrum models. Presented here is a proteomics search engine that offers a new variation of the standard approach, with improved results. The proposed method employs information theory and probabilistic information retrieval on a pre-computed and indexed fragmentation database generating a peptide-to-spectrum match (PSM) score modeled on fragment ion frequency. As a result, the direct application of modern document mining, allows for treating the collection of peptides as a corpus and corresponding fragment ions as indexable words, leveraging ready-built search engines and common predefined ranking algorithms. Fast and accurate PSM matches are achieved yielding a 5-10% higher rate of peptide identities than current database mining methods. Immediate applications of this search engine are aimed at identifying peptides from large sequence databases consisting of homologous proteins with minor sequence variations, such as genetic variation expected in the human population.Competing Interest StatementThis research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Researchers have commercial interests in applications of the technology described herein.View Full Text
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jones, J.. 2020-02-07. A Pre-computed Probabilistic Molecular Search Engine for Tandem Mass Spectrometry Proteomics. https://doi.org/10.1101/2020.02.06.937870
Cite the original work for its findings. Save a collection to share your selection of sources.