PRReMS: Parallel Regularised Regression Model Search for bio-signature discovery
There is increasing interest in developing point of care tests to diagnose disease and predict prognosis based upon biomarker signatures of RNA or protein expression levels. Technology to measure the required biomarkers accurately and in a time-frame useful to health care professionals will be easier to develop by minimising the number of biomarkers measured. In this paper we describe the Parallel Regularised Regression Model Search (PReMS) method which is designed to estimate parsimonious prediction models. Given a set of potential biomarkers PReMS searches over many logistic regression models constructed from optimal subsets of the biomarkers, iteratively increasing the model size. Zero centred Gaussian prior distributions are assigned to all regression coefficients to induce shrinkage. The method estimates the optimal shrinkage parameter, optimal model for each model size and the optimal model size. We apply PReMS to six freely available data sets and compare its performance with the LASSO and SCAD algorithms in terms of the number of covariates in the model, model accuracy, as measured by the area under the receiver operator curve (AUC) and root predicted mean square error, and model calibration. We show that PReMS typically selects models with fewer biomarkers than both the LASSO and SCAD algorithms but has comparable predictive accuracy.\n\nAvailability: (PReMS) is freely available as an R package https://github.com/clivehoggart/PReMS