Search bioRxiv⌕ Search

Biology subjects

Goudo, M.

Publications and source records attributed to Goudo, M..

2 recordsLinked to original sources

Comparison of classification accuracy and feature selection between sparse and non-sparse modeling of metabolomics data

Machine learnings such as multivariate analyses and clustering have been frequently used for metabolomics data analyses. In metabolomics data analyses, how much difference there is between the results calculated by supervised and unsupervised learning models is an interesting topic. Since metabolomics data include hundreds to thousands of metabolites greater than the sample numbers, only a small fraction of metabolites is relevant to the phenotype of interest. For this reason, sparse mechanisms have been introduced into many machine learning models. However, its explanatory power decreases when the number of explanatory variables is reduced to an extreme level. In this paper, serum lipidomic data of breast cancer patients (1) pre/post-menopause and (2) before/after neoadjuvant chemotherapy was chosen as one of metabolomics data. Here, this data was analyzed by partial least squares (PLS) for regression and K-means and hierarchical clustering for clustering. Results were also compare with the sparse modeling. Between the non-sparse and sparse modeling accuracy, there is no significant difference. Metabolite subsets selected by sparse modeling were almost identical to the PLS-selected features. At the same time, several metabolites were consistently selected regardless of the algorithm used. These results contribute to exploring biomarkers in high-dimensional metabolomics datasets.

bioinformatics↗

The usefulness of sparse k-means in metabolomics data: Anexample from breast cancer data

In processing metabolomics data, multidimensional quantitative data from thousands of metabolites are often sparse, that is, only a small fraction of metabolites are relevant to the phenotype of interest. Clustering is therefore used to discover subtypes from omics data. Sparse processing, which selects important metabolites from the total omics data, is an effective clustering technique. This study investigated the effectiveness of sparse k-means for metabolomics data. Specifically, sparse k-means was used to cluster blood lipid metabolite data of breast cancer patients in two studies: (1) before and after menopause, and (2) pre- and postoperative chemotherapy. In both cases, sparse k-means showed comparable discrimination accuracy with fewer metabolites than k-means. Furthermore, when the L1 norm values were varied, no significant changes were observed. The mean silhouette coefficients of sparse k-means and k-means were (1) 0.38 {+/-} 0.14 (S.D.) and 0.17 {+/-} 0.01, (2) 0.38 {+/-} 0.07 and 0.17 {+/-} 0.01, indicating that feature selection using sparse k-means can improve clustering results. In addition, metabolite selection using sparse k-means was consistent regardless of the test data or the constrained value of the L1 norm, indicating robustness.

bioinformatics↗