Search bioRxiv⌕ Search

Biology subjects

Urazbakhtin, S.

Publications and source records attributed to Urazbakhtin, S..

2 recordsLinked to original sources

Isobaric labeling update in MaxQuant

We present an update of the MaxQuant software for isobaric labeling data and evaluate its performance on benchmark datasets. Impurity correction factors can be applied to labels mixing C- and N-type reporter ions, such as TMT Pro. Application to a single-cell multi-species mixture benchmark shows high accuracy of the impurity-corrected results. TMT data recorded with FAIMS separation can be analyzed directly in MaxQuant without splitting the raw data into separate files per FAIMS voltage. Weighted median normalization, is applied to several datasets, including large-scale human body atlas data. In the benchmark datasets the weighted median normalization either removes or strongly reduces the batch effects between different TMT plexes and results in clustering by biology. In datasets including reference channels, we find that weighted median normalization performs as well or better when the reference channels are ignored and only the sample channel intensities are used, suggesting that the measurement of reference channels is unnecessary when using weighted median normalization in MaxQuant. We demonstrate that MaxQuant including the weighted median normalization performs well on multi-notch MS3 data, as well as on phosphorylation data. MaxQuant is freely available for any purpose and can be downloaded from https://www.maxquant.org/.

bioinformatics↗

Tesorai Search: Large pretrained model boosts identifications in mass spectrometry proteomics without the need for Percolator.

The original mass spectrometry search engines used simple algorithms for peptide identification. Recent tools improved accuracy by adding several extra components such as fragment ion intensities or retention times prediction and training target-decoy classifiers on-the-fly, leading to sometimes inconsistent results. Our study explores the impact of replacing those extra components with a deep-learning pretrained model that directly learns the complex relationship between the full spectra and associated peptide sequence, without using decoys. This simplified workflow has fewer parameters to tweak, making it easier to use and perform robustly on data from instruments and use-cases never seen during training. Surprisingly, our approach consistently identifies more peptides than FragPipe, PEAKS, and Proteome Discoverer (12%, 9%, and 21% more, respectively, across a range of datasets). Tesorai Search is also fast - 250 immunopeptidomics searches in 45 minutes - and free for academics, available as a webserver at console.tesorai.com.

bioinformatics↗