bioRxiv · 10.1101/2022.09.19.508550
Interpretable machine learning with tree-based Shapley additive explanations: application to metabolomics datasets for binary classification
Abstract
Machine learning (ML) models are used in clinical metabolomics studies most notably for biomarker discoveries, to identify metabolites that discriminate between a case and control group. To improve understanding of the underlying biomedical problem and to bolster confidence in these discoveries, model interpretability is germane. In metabolomics, partial least square discriminant analysis (PLS-DA) and its variants are widely used, partly due to the models interpretability with the Variable Influence in Projection (VIP) scores, a global interpretable method. Herein, Tree-based Shapley Additive explanations (SHAP), an interpretable ML method grounded in game theory, was used to explain ML models with local explanation properties. In this study, ML experiments (binary classification) were conducted for three published metabolomics datasets using PLS-DA, random forests, gradient boosting, and extreme gradient boosting (XGBoost). Using one of the datasets, PLS-DA model was explained using VIP scores, while a tree-based model was interpreted using Tree SHAP. The results show that SHAP has a more explanation depth than PLS-DAs VIP, making it a powerful method for rationalizing machine learning predictions from metabolomics studies.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bifarin, O. O.. 2022-09-19. Interpretable machine learning with tree-based Shapley additive explanations: application to metabolomics datasets for binary classification. https://doi.org/10.1101/2022.09.19.508550
Cite the original work for its findings. Save a collection to share your selection of sources.