bioRxiv · 10.1101/2023.05.31.542998
Quality control and annotation of variant peptides identified through proteogenomics
Abstract
Variant peptides resulting from translation of single nucleotide polymorphisms (SNPs) can lead to aberrant or altered protein functions and thus hold translational potential for disease diagnosis, therapeutics and personalized medicine. Variant peptides detected by proteogenomics are fraught with high number of false positives. Class-specific FDR along with ad-hoc post-search filters have been employed to tackle this issue, but there is no uniform and comprehensive approach to assess variant quality. These protocols are mostly manual or tedious, and not accessible across labs. We present a software tool, PgxSAVy, for the quality control of variant peptides. PgxSAVy provides a rigorous framework for quality control and annotations of variant peptides on the basis of (i) variant quality, (ii) isobaric masses, and (iii) disease annotation. PgxSAVy was able to segregate true and false variants with 98.43% accuracy on simulated data. We then used [~]2.8 million spectra (PXD004010 and PXD001468) and identified 12,705 variant PSMs, of which PgxSAVy evaluated 3028 (23.8%), 1409 (11.1%) and 8268 (65.1%) as confident, semi-confident and doubtful respectively. PgxSAVy also annotates the variants based on their pathogenicity and provides support for assisted manual validation. In these datasets, it identified previously found variants as well some novel variants not seen in original studies. The confident variants identified the importance of mutations in glycolysis and gluconeogenesis pathways in Alzheimers disease. The analysis of proteins carrying variants can provide fine granularity in discovering important pathways. PgxSAVy will advance personalized medicine by providing a comprehensive framework for quality control and prioritization of proteogenomics variants. AvailabilityPgxSAVy is freely available at https://github.com/anuragraj/PgxSAVy Key PointsO_LIVariant peptide in proteogenomics have high rates of false positives C_LIO_LIclass-specific FDR is not sufficiently effective, and tedious manual filtering is not scalable C_LIO_LIWe developed PgxSAVy for automated quality control and disease annotation of variant peptides from proteogenomics search results C_LIO_LIPgxSAVy was validated using simulation data and manually annotated variant PSMs C_LIO_LIIndependent application on large datasets on Alzheimers and HEK cell lines demonstrated that PgxSAVy discovered known and novel mutations with important biological roles. C_LI Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=161 HEIGHT=200 SRC="FIGDIR/small/542998v2_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@187a515org.highwire.dtl.DTLVardef@67471corg.highwire.dtl.DTLVardef@6d90e9org.highwire.dtl.DTLVardef@144c6c1_HPS_FORMAT_FIGEXP M_FIG C_FIG
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Raj, A., Aggarwal, S., Yadav, A. K., Dash, D.. 2023-06-03. Quality control and annotation of variant peptides identified through proteogenomics. https://doi.org/10.1101/2023.05.31.542998
Cite the original work for its findings. Save a collection to share your selection of sources.