bioRxiv · 10.1101/754630
Revisiting Feature Selection with Data Complexity for Biomedicine
Abstract
The identification of biomarkers or predictive features that are indicative of a specific biological or disease state is a major research topic in biomedical applications. Several feature selection(FS) methods ranging from simple univariate methods to recent deep-learning methods have been proposed to select a minimal set of the most predictive features. However, there still lacks the answer to the question of "which method to use when". In this paper, we study the performance of feature selection methods with respect to the underlying datasets statistics and their data complexity measures. We perform a comparative study of 11 feature selection methods over 27 publicly available datasets evaluated over a range of number of selected features using classification as the downstream task. We take the first step towards understanding the FS methods performance from the viewpoint of data complexity. Specifically, we (empirically) show that as regard to classification, the performance of all studied feature selection methods is highly correlated with the error rate of a nearest neighbor based classifier. We also argue about the non-suitability of studied complexity measures to determine the optimal number of relevant features. While looking closely at several other aspects, we also provide recommendations for choosing a particular FS method for a given dataset.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dong, T. N., Winkler, L., Khosla, M.. 2019-09-07. Revisiting Feature Selection with Data Complexity for Biomedicine. https://doi.org/10.1101/754630
Cite the original work for its findings. Save a collection to share your selection of sources.