A Statistical Guideline for Analyzing Repeated-Measures Proteomics Data
Mass spectrometry-based proteomics increasingly supports study designs with repeated measurements, such as longitudinal sampling and spatial tissue profiling. These designs offer rich biological insights but introduce statistical challenges due to within-subject correlations and variance heterogeneity across time points or anatomical locations. Despite the growing use, best practices for analyzing repeated-measures proteomics data remain limited. We systematically compare commonly applied methods, including paired t-tests and linear mixed-effects models with correlation induced by random effects, to more flexible covariance-pattern models with unstructured covariance. Through simulation studies, we evaluate false-positive rates and statistical power under varying conditions of covariance heterogeneity. Our results demonstrate substantial limitations of conventional approaches and highlight the robustness and superior performance of unstructured covariance pattern models for repeated measures proteomics. This approach accommodates diverse study designs supporting accurate inference and reproducibility. We demonstrate that conventional F-tests can be severely biased, not just for small sample sizes but also moderately large ones, while using pairwise t-tests or applying parametric bootstrapping to evaluate the p-value of the F-tests mitigates the problem. We provide real world data examples showing how to perform valid and informative statistical analyses for temporal and spatial proteomics studies with repeated measurements.