Search bioRxiv⌕ Search

Biology subjects

Bourassa, F.

Publications and source records attributed to Bourassa, F..

3 recordsLinked to original sources

Discovery of non-canonical proteins through modification-aware proteogenomics

ShortThe SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. LongO_ST_ABSBackgroundC_ST_ABSThe SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. ResultsCompared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. ConclusionsIn this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.

molecular biology↗

Detection of human unannotated microproteins by mass spectrometry-based proteomics: a community assessment

Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)- based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.

genomics↗

A Proximity MAP of RAB GTPases

RAB GTPases are the most abundant family of small GTPases and regulate multiple aspects of membrane trafficking events, from cargo sorting to vesicle budding, transport, docking, and fusion. To regulate these processes, RABs are tightly regulated by guanine exchange factors (GEFs) and GTPase-activating proteins (GAPs). Activated RABs recruit effector proteins that regulate trafficking. Identifying RAB-associated proteins has proven to be difficult because their association with interacting proteins is often transient. Recent advances in proximity labeling approaches that allow for the covalent labeling of neighbors of proteins of interest now permit the cataloging of proteins in the vicinity of RAB GTPases. Here, we report APEX2 proximity labeling of 23 human RABs and their neighboring proteomes. We have used bioinformatic analyses to map specific proximal proteins for an extensive array of RAB GTPases, and RAB localization can be inferred from their adjacent proteins. Focusing on specific examples, we identified a physical interaction between RAB25 and DENND6A, which affects cell migration. We also show functional relationships between RAB14 and the EARP complex, or between RAB14 and SHIP164 and its close ortholog UHRF1BP1. Our dataset provides an extensive resource to the community and helps define novel functional connections between RAB GTPases and their neighboring proteins.

cell biology↗