Search bioRxiv⌕ Search

Biology subjects

van Puyenbroeck, S.

Publications and source records attributed to van Puyenbroeck, S..

3 recordsLinked to original sources

Limitations of de novo sequencing in resolving sequence ambiguity

De novo peptide sequencing enables peptide identification from fragmentation spectra without relying on sequence databases. However, incomplete spectra create ambiguity, making unambiguous identification challenging. Recent deep learning advances have produced numerous de novo models that predict sequences and refine peptide-spectrum matches under such conditions. Yet, their relative strengths, weaknesses, and ability to handle spectrum ambiguity remain unclear. Here, we benchmark eight state-of-the-art models on three publicly available proteomics datasets, comparing performance using established metrics and quantifying inter-model agreement. We assess post-processing approaches, including iterative refinement, rescoring, and reranking, for their ability to improve identification accuracy, and perform an error analysis to identify common mispredictions and their causes. Model performance varied, with considerable overlap of correct identifications. Post-processing yielded no or only modest improvements. Most sequencing errors were model-independent and driven by limited fragment ion coverage, a limitation also observed in database searches with large search spaces.

bioinformatics↗

MLMarker: A machine learning framework for tissue inference and biomarker discovery

MLMarker is a machine learning tool that computes continuous tissue similarity scores for proteomics data, addressing the challenge of interpreting complex or sparse datasets. Trained on 34 healthy tissues, its Random Forest model generates probabilistic predictions with SHAP-based protein-level explanations. A penalty factor corrects for missing proteins, improving robustness for low-coverage samples. Across three public datasets, MLMarker revealed brain-like signatures in cerebral melanoma metastases, achieved high accuracy in a pan-cancer cohort, and identified brain and pituitary origins in biofluids. MLMarker provides an interpretable framework for tissue inference and hypothesis generation, available as a Python package and Streamlit app.

bioinformatics↗

InstaNovo-P: A de novo peptide sequencing model for phosphoproteomics

Phosphorylation, a crucial post-translational modification (PTM), plays a central role in cellular signaling and disease mechanisms. Mass spectrometry-based phosphoproteomics is widely used for system-wide characterization of phosphorylation events. However, traditional methods struggle with accurate phosphorylated site localization, complex search spaces, and detecting sequences outside the reference database. Advances in de novo peptide sequencing offer opportunities to address these limitations, but have yet to become integrated and adapted for phosphoproteomics datasets. Here, we present InstaNovo-P, a phosphorylation specific version of our transformer-based InstaNovo model, fine-tuned on extensive phosphoproteomics datasets. InstaNovo-P significantly surpasses existing methods in phosphorylated peptide detection and phosphorylated site localization accuracy across multiple datasets, including complex experimental scenarios. Our model robustly identifies peptides with single and multiple phosphorylated sites, effectively localizing phosphorylation events on serine, threonine, and tyrosine residues. We experimentally validate our model predictions by studying FGFR2 signaling, further demonstrating that InstaNovo-P uncovers phosphorylated sites previously missed by traditional database searches. These predictions align with critical biological processes, confirming the models capacity to yield valuable biological insights. InstaNovo-P adds value to phosphoproteomics experiments by effectively identifying biologically relevant phosphorylation events without prior information, providing a powerful analytical tool for the dissection of signaling pathways.

bioinformatics↗