Benchmark of lasso-like penalties in the Cox model for TCGA datasets reveal improved performance with pre-filtering and wide differences between cancers
MotivationPrediction of patient survival from tumor molecular omics data is a key step toward personalized medicine. With this aim, the databases available are growing, with the collection of various omics characterizations of patient tumors, together with their associated clinical outcomes for weeks to years of follow-up. Cox models with variable selection used with RNA profiling datasets are popular for identification of prognostic biomarkers and for clinical predictions. However, these models are confronted with the curse of dimensionality, as the number p of covariates (genes) can greatly exceed the number n of patients. To tackle this problem, variance-based pre-filtering and penalization methods are popular for dimension reduction. In the present paper, we study the impact of a pre-filtering step based on gene variability, and we evaluate the performance of the lasso penalization of the Cox model and four variants (i.e., elastic net, adaptive elastic net, ridge, univariate Cox) in terms of prediction, selection and stability. ResultsFirst, we show that the prediction capacity with the Cox penalties method is cancer dependent. Second, we develop a methodology to fix a threshold to filter out genes with low variability without losing prediction capacity. Third, we show that it is best not to use the Cox model to select prognostic biomarkers, as its false discovery proportion is always [≥] 50%. Finally, to predict overall survival, we can suggest the use of the ridge penalty, or the elastic net if a more parsimonious model is needed, after the pre-filtering step. AvailabilityWe provide the R script generated to reproduce all of the figures presented in this article. Supplementary informationSupplementary Figures and R scripts are available.