Search bioRxiv⌕ Search

Biology subjects

Penanes, P. A.

Publications and source records attributed to Penanes, P. A..

2 recordsLinked to original sources

DeepLC introduces transfer learning for accurate LC retention time prediction and adaptation to substantially different modifications and setups

While LC retention time prediction of peptides and their modifications has proven useful, widespread adoption and optimal performance are hindered by variations in experimental parameters. These variations can render retention time prediction models inaccurate and dramatically reduce the value of predictions for identification, validation, and DIA spectral library generation. To date, mitigation of these issues has been attempted through calibration or by training bespoke models for specific experimental setups, with only partial success. We here demonstrate that transfer learning can successfully overcome these limitations by leveraging pre-trained model parameters. Remarkably, this approach can even fit highly performant models to substantially different peptide modifications and LC conditions than those on which the model was originally trained. This impressive adaptability of transfer learning makes it a highly robust solution for accurate peptide retention time prediction across a very wide variety of imaginable proteomics workflows.

bioinformatics↗

Potential of Negative Ion Mode Proteomics: MS1-Only Approach

Current proteomics approaches rely almost exclusively on using positive ionization mode, which results in inefficient ionization of many acidic peptides. With an equal quantity of acidic and basic proteins and, correspondingly, the similar number for their derived peptides in case of the human proteome, this inefficient ionization poses both a substantial challenge and a potential. In this work, we study the efficiency of protein identification in the bottom-up proteomic analysis performed in negative ionization mode, using the recently introduced MS1-only ultra-fast data acquisition method DirectMS1. This method is based on accurate peptide mass measurements and predicted retention times. Our method achieves the highest rate of protein identifications in negative ion mode to date, with over 1,000 proteins identified in a human cell line at a 1% false discovery rate using a single-shot 10-min separation gradient, which is comparable with hours-long MS/MS-based analyses. Evaluating the proteins as a function of pI indicated preferable identification of the acidic part of the proteome. Optimization of separation and mass spectrometric experimental conditions facilitated the performance of the method with the best results in terms of spray stability and signal abundance obtained using mobile buffers at 2.5 mM imidazole and 3% isopropanol. The work also highlighted the complementarity of data acquired in positive and negative modes: Combining the results for all replicates for both polarities, the number of identified proteins increased up to 1,774. Finally, we performed analysis of the methods efficiency when different proteases are used for protein digestion. Among the four studied proteases (LysC, GluC, AspN, and trypsin), we found that trypsin and LysC performed best in terms of protein identification yield. Thus, digestion procedures used for positive mode proteomics can be efficiently utilized for analysis in negative ion mode.

molecular biology↗