Fast-Part: Fast and Accurate Data Partitioning for Biological Sequence Analysis
Developing effective machine learning models for classifications of biological sequences depends heavily on the quality of the training and test datasets split. Existing tools are either computationally expensive, unable to maintain the desired level of similarity between the training and test datasets, or unable to retain training-test ratio stratification. Here, we present Fast-Part, a fast and accurate sequence data partitioning tool that ensures strict homology separation between the training and test datasets and the best possible training: test stratification ratio, and at the same time, is computationally fast. Fast-Part demonstrates rapid and accurate partitioning performance across diverse protein sequence datasets and maintains strict partitioning compared to the existing tools. Fast-Part can handle massive datasets and maintain strict homology partitioning.