bioRxiv · 10.64898/2025.12.12.694008
ProteoForge: An Imputation-Aware Framework for Differential Proteoform Discovery in Bottom-Up Proteomics
Abstract
The human genome contains approximately 20,000 protein-coding genes. However, millions of diverse protein variants, called proteoforms, exist. Despite originating from the same gene, proteoforms often have distinct biological roles. In bottom-up proteomics, the aggregation of peptide measurements into protein-level quantities often obscures this information. Existing methods for proteoform deconvolution are limited by their handling of missing data, which can introduce significant bias. To address this we developed ProteoForge, which builds on an imputation-aware statistical model to identify and group co-varying peptides into quantitatively differential proteoforms (dPFs). Benchmarking against existing deconvolution methods demonstrated that ProteoForge provides high accuracy and stability in datasets with high rates of missing values, complex experimental designs, or varying signal strengths. Application of ProteoForge to proteomics data from lung cancer cells under hypoxia revealed extensive proteoform-level regulation hidden by standard protein-level analysis.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ergin, E. K., Conrrero, A., Ferguson, K. M., Lange, P. F.. 2025-12-16. ProteoForge: An Imputation-Aware Framework for Differential Proteoform Discovery in Bottom-Up Proteomics. https://doi.org/10.64898/2025.12.12.694008
Cite the original work for its findings. Save a collection to share your selection of sources.