Search bioRxivSearch

Biology subjects

Eng, J. K.

Publications and source records attributed to Eng, J. K..

2 recordsLinked to original sources

Full-featured, real-time database searching platform enables fast and accurate multiplexed quantitative proteomics.

Multiplexed quantitative analyses of complex proteomes enable deep biological insight. While a multitude of workflows have been developed for multiplexed analyses, the most quantitatively accurate method (SPS-MS3) suffers from long acquisition duty cycles. We built a new, real-time database search (RTS) platform, Orbiter, to combat the SPS-MS3 methods longer duty cycles. RTS with Orbiter enables the elimination of SPS-MS3 scans if no peptide matches to a given spectrum. With Orbiters online proteomic analytical pipeline, which includes RTS and false discovery rate analysis, it was possible to process a single spectrum database search in less than 10 milliseconds. The result is a fast, functional means to identify peptide spectral matches using Comet, filter these matches, and more efficiently quantify proteins of interest. Importantly, the use of Comet for peptide spectral matching allowed for a fully featured search, including analysis of post-translational modifications, with well-known and extensively validated scoring. These data could then be used to trigger subsequent scans in an adaptive and flexible manner. In this work we tested the utility of this adaptive data acquisition platform to improve the efficiency and accuracy of multiplexed quantitative experiments. We found that RTS enabled a 2-fold increase in mass spectrometric data acquisition efficiency. Orbiters RTS was able to quantify more than 8000 proteins across 10 proteomes in half the time of an SPS-MS3 analysis (18 hours for RTS, 36 hours for SPS-MS3).

cell biology

Proteomics Standards Initiative Extended FASTA Format (PEFF)

Mass spectrometry-based proteomics enables the high-throughput identification and quantification of proteins, including sequence variants and post-translational modifications (PTMs), in biological samples. However, most workflows require that such variations be included in the search space used to analyze the data, and doing so remains challenging with most analysis tools. In order to facilitate the search for known sequence variants and PTMs, the Proteomics Standards Initiative (PSI) has designed and implemented the PSI Extended FASTA Format (PEFF). PEFF is based on the very popular FASTA format but adds a uniform mechanism for encoding substantially more metadata about the sequence collection as well as individual entries, including support for encoding known sequence variants, PTMs, and proteoforms. The format is very nearly backwards compatible, and as such, existing FASTA parsers will require little or no changes to be able to read PEFF files as FASTA files, although without supporting any of the extra capabilities of PEFF. PEFF is defined by a full specification document, controlled vocabulary terms, a set of example files, software libraries, and a file validator. Popular software and resources are starting to support PEFF, including the sequence search engine Comet and the knowledge bases neXtProt and UniProtKB. Widespread implementation of PEFF is expected to further enable proteogenomics and top-down proteomics applications by providing a standardized mechanism for encoding protein sequences and their known variations. All the related documentation, including the detailed file format specification and example files, are available at http://www.psidev.info/peff.

bioinformatics