Search bioRxivSearch

Biology subjects

Egert, J.

Publications and source records attributed to Egert, J..

1 recordsLinked to original sources

DIMA: Data-driven selection of a suitable imputation algorithm

MotivationImputation is a prominent strategy when dealing with missing values (MVs) in proteomics data analysis pipelines. However, the performance of different imputation methods is difficult to assess and varies strongly depending on data characteristics. To overcome this issue, we present the concept of a data-driven selection of a suitable imputation algorithm (DIMA). ResultsThe performance and broad applicability of DIMA is demonstrated on 121 quantitative proteomics data sets from the PRIDE database and on simulated data consisting of 5 - 50% MVs with different proportions of missing not at random and missing completely at random values. DIMA reliably suggests a high-performing imputation algorithm which is always among the three best algorithms and results in a root mean square error difference ({Delta}RMSE) [≤] 10% in 84% of the cases. Availability and ImplementationSource code is freely available for download at github.com/clemenskreutz/OmicsData.

bioinformatics