Search bioRxiv⌕ Search

Biology subjects

Devreese, R.

Publications and source records attributed to Devreese, R..

8 recordsLinked to original sources

Predicting and Elucidating Peptide Retention Mechanisms with Graph Attention Networks

Liquid chromatography (LC) is a key technology in bottom-up proteomics, separating proteolytic peptides to decrease sample complexity, enhance coverage, and increase the robustness of protein identification and quantification. Although high-resolution mass spectrometry has advanced significantly, comparable progress in LC has lagged, primarily due to a limited understanding of peptide-column interactions. To bridge this knowledge gap, we introduce a novel deep learning model (PeptideGNN) based on a Graph Neural Network (GNN) architecture to model and elucidate peptide behaviors across various separation conditions. Trained to accurately predict peptide retention times on ten diverse proteomic datasets, the model subsequently employed a saliency mapping technique to interpret the underlying retention mechanisms. Our model consistently outperformed existing retention-time predictors across multiple datasets, while the saliency mapping, importantly, revealed insights into peptide-stationary phase interactions, highlighting the effects of neighboring amino acids, post-translational modifications (PTMs), chromato-graphic columns, and mobile phase additives on peptide retention.

bioinformatics↗

LFQ Benchmark Dataset - Generation Beta: Assessing Modern Proteomics Instruments and Acquisition Workflows with High-Throughput LC Gradients

Recent advances in liquid chromatography-mass spectrometry (LC-MS) have accelerated the adoption of high-throughput workflows that deliver deep proteome coverage using minimal sample amounts. This trend is largely driven by clinical and single-cell proteomics, where sensitivity and reproducibility are essential. Here, we extend our previous benchmark dataset (PXD028735) using next-generation LC-MS platforms optimized for rapid proteome analysis. We generated an extensive DDA/DIA dataset using a human-yeast-E. coli hybrid proteome. The proteome sample was distributed across multiple laboratories together with standardized analytical protocols specifying two short LC gradients (5 and 15 min) and low sample input amounts. This dataset includes data acquired on four different platforms, and features new scanning quadrupole-based implementations, extending coverage across different instruments and acquisition strategies. Our comprehensive evaluation highlights how technological advances and reduced LC gradients may affect proteome depth, quantitative precision, and cross-instrument consistency. The release of this benchmark dataset via ProteomeXchange (PXD070049 and PXD071205), allows for the acceleration of cross-platform algorithm development, enhance data mining strategies, and supports standardization of short-gradient, high-throughput LC-MS-based proteomics.

bioinformatics↗

ProteoBench: the community-curated platform for comparing proteomics data analysis workflows

Mass spectrometry (MS)-based proteomics is a well-established strategy for analyzing complex biological mixtures. Many MS instruments and data acquisition strategies are available, and the data they acquire differ substantially, thus requiring tailored analysis algorithms. Hence, many dedicated bioinformatics workflows are developed. These are in constant evolution, and the community lacks a centralized platform for comparing their performance. Here, we propose ProteoBench, a single platform that brings together software developers and software users to provide an ever-evolving comparison of state-of-the-art proteomics data processing tools. ProteoBench is an open-source resource that enables the community to evaluate data analysis workflows, develop benchmarking modules dedicated to specific comparisons, and discuss the best methods to compare software tools. The platform ensures that the benchmark evolves alongside advances in proteomics data analysis workflows. ProteoBench guides researchers towards the best-suited tool and parameters for their specific project and data according to their needs, and developers can test their newly developed tools or workflows privately, before adding them as public references. This community-driven effort will increase transparency and reproducibility between MS data analysis workflows, as well as facilitate the development and publication of software workflows in the field.

bioinformatics↗

iDeepLC: chemical structure information yields improved retention time prediction of peptides with unseen modifications

Deep learning has notably advanced the field of liquid chromatography-mass spectrometry-based proteomics. Accurate prediction of peptide retention times significantly enhances our ability to match LC-MS data with the correct peptides and proteins, especially for DIA data. While numerous models predict peptide LC retention times with high accuracy, few can accurately predict the retention times of chemically modified peptides, particularly those with modifications not encountered during model training. In our previously developed DeepLC model, accurate predictions could be made for unseen modifications by leveraging the chemical composition of (modified) residues. Here, however, we present a further enhancement of this model based on chemical structural information. The resulting model, called iDeepLC, shows overall more accurate predictions, and better generalization performance for predicting the retention time of unseen modifications than DeepLC. iDeepLC is freely available as open-source software under the Apache2 license and can be found at https://github.com/CompOmics/iDeepLC.

bioinformatics↗

DeepLC introduces transfer learning for accurate LC retention time prediction and adaptation to substantially different modifications and setups

While LC retention time prediction of peptides and their modifications has proven useful, widespread adoption and optimal performance are hindered by variations in experimental parameters. These variations can render retention time prediction models inaccurate and dramatically reduce the value of predictions for identification, validation, and DIA spectral library generation. To date, mitigation of these issues has been attempted through calibration or by training bespoke models for specific experimental setups, with only partial success. We here demonstrate that transfer learning can successfully overcome these limitations by leveraging pre-trained model parameters. Remarkably, this approach can even fit highly performant models to substantially different peptide modifications and LC conditions than those on which the model was originally trained. This impressive adaptability of transfer learning makes it a highly robust solution for accurate peptide retention time prediction across a very wide variety of imaginable proteomics workflows.

bioinformatics↗

Collisional cross-section prediction for multiconformational peptide ions with IM2Deep

Peptide collisional cross-section (CCS) prediction is complicated by the tendency of peptide ions to exhibit multiple conformations in the gas phase. This adds further complexity to downstream analysis of proteomics data, for example for identification or quantification through feature finding. Here, we present an improved version of IM2Deep that is trained on a carefully curated dataset to predict CCS values of multiconformational peptides. The training data is derived from a large and comprehensive set of publicly available datasets. This comprehensive training dataset together with a tailored architecture allows for the accurate CCS prediction of multiple peptide conformational states. Furthermore, the enhanced IM2Deep model also retains high precision for peptides with a single observed conformation. IM2Deep is publicly available under a permissive open source license at https://github.com/compomics/IM2Deep.

bioinformatics↗

Maximizing immunopeptidomics-based bacterial epitope discovery by multiple search engines and rescoring

Mass spectrometry-based discovery of bacterial immunopeptides presented by infected cells allows untargeted discovery of bacterial antigens that can serve as vaccine candidates. However, reliable identification of bacterial epitopes is challenged by their extreme low abundance. Here, we describe an optimized bioinformatical framework to enhance the confident identification of bacterial immunopeptides. Immunopeptidomics data of cell cultures infected with Listeria monocytogenes were searched by four different search engines, PEAKS, Comet, Sage and MSFragger, followed by data-driven rescoring with MS2Rescore. Compared to individual search engine results, this integrated workflow boosted immunopeptide identification by an average of 27% and led to the high-confidence detection of 18 additional bacterial peptides (+27%) matching 15 different Listeria proteins (+36%). Despite the strong agreement between the search engines, a small number of spectra (< 1%) had ambiguous matches to multiple peptides and were excluded to ensure high-confident identifications. Finally, we demonstrate our workflow with sensitive timsTOF SCP data acquisition and find that rescoring, now with inclusion of ion mobility features, identifies 76% more peptides compared to Q Exactive HF acquisition. Together, our results demonstrate how integration of multiple search engine results along with data-driven rescoring maximizes immunopeptide identification, boosting the detection of high-confidence bacterial epitopes for vaccine development.

molecular biology↗

TIMS2Rescore: A DDA-PASEF optimized data-driven rescoring pipeline based on MS2Rescore

The high throughput analysis of proteins with mass spectrometry (MS) is highly valuable for understanding human biology, discovering disease biomarkers, identifying therapeutic targets, and exploring pathogen interactions. To achieve these goals, specialized proteomics subfields - such as plasma proteomics, immunopeptidomics, and metaproteomics - must tackle specific analytical challenges, such as an increased identification ambiguity compared to routine proteomics experiments. Technical advancements in MS instrumentation can counter these issues by acquiring more discerning information at higher sensitivity levels, as is exemplified by the incorporation of ion mobility and parallel accumulation - serial fragmentation (PASEF) technologies in timsTOF instruments. In addition, AI-based bioinformatics solutions can help overcome ambiguity issues by integrating more data into the identification workflow. Here, we introduce TIMS2Rescore, a data-driven rescoring workflow optimized for DDA-PASEF data from timsTOF instruments. This platform includes new timsTOF MS2PIP spectrum prediction models and IM2Deep, a new deep learning-based peptide ion mobility predictor. Furthermore, to fully streamline data throughput, TIMS2Rescore directly accepts Bruker raw mass spectrometry data, and search results from ProteoScape and many other search engines, including MS Amanda and PEAKS. We showcase TIMS2Rescore performance on plasma proteomics, immunopeptidomics (HLA class I and II), and metaproteomics data sets. TIMS2Rescore is open-source and freely available at https://github.com/compomics/tims2rescore.

bioinformatics↗