Crimp: fast and scalable cluster relabeling based on impurity minimization
MotivationTo analyze population structure based on multilocus geno-type data, a variety of popular tools perform model-based clustering, as-signing individuals to a prespecified number of ancestral populations. Since such methods often involve stochastic components, it is a common practice to perform multiple replicate analyses based on the same input data and parameter settings. Their results are typically affected by the label-switching phenomenon, which complicates their comparison and summary. Available tools allow to mitigate this problem, but leave room for improvements, in particular, regarding large input datasets. ResultsIn this work, I present CO_SCPLOWRIMPC_SCPLOW, a lightweight command-line tool, which offers a relatively fast and scalable heuristic to align clusters across multiple replicate clusterings consisting of the same number of clusters. For small problem sizes, an exact algorithm can be used as alternative. Additional features include row-specific weights, input and output files similar to those of CLUMPP (Jakobsson & Rosenberg, 2007), and the evaluation of a given solution in terms of either CLUMPPs and its own objective functions. Benchmark analyses show that CO_SCPLOWRIMPC_SCPLOW, especially when applied to larger datasets, tends to outperform alternative tools considering runtime requirements and various quality measures. AvailabilityCO_SCPLOWRIMPC_SCPLOWs source code along with precompiled binaries for Linux and Windows, usage guidelines and benchmark code are freely available at https://github.com/ulilautenschlager/crimp. Contactulrich.lautenschlager@ur.de