bioRxiv · 10.1101/2025.01.22.634321
Detecting and avoiding homology-based data leakage in genome-trained sequence models
Abstract
Models that predict function from DNA sequence have become critical tools in deciphering the roles of genomic sequences and genetic variation within them. However, traditional approaches for dividing the genomic sequences into training data, used to create the model, and test data, used to determine the models performance on unseen data, fail to account for the widespread homology within genomes. Using simulations, we illustrate how homology-based data leakage can lead to overestimation of model performance. Across a variety of genomics models, we demonstrate that performance on test sequences varies systematically by their similarity with training sequences. Models generally perform well on distant sequences, reflecting the application of learned generalizable principles. At higher and intermediate similarity, models rely on memorized associations, inflating performance when function is conserved between homologs but failing when homologous sequences have functionally diverged. To dissect and mitigate these effects, we introduce hashFrag, a scalable solution for homology detection and data partitioning. Using hashFrag, we demonstrate how to create homology-aware evaluations of model performance, and improve model generalizability by providing improved splits for model training. Altogether, we establish how homology creates a systematic bias in genome-trained models and must be accounted for to ensure reliable evaluation of sequence-to-function predictors.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rafi, A. M., Kiyota, B., Yachie, N., de Boer, C. G.. 2025-01-24. Detecting and avoiding homology-based data leakage in genome-trained sequence models. https://doi.org/10.1101/2025.01.22.634321
Cite the original work for its findings. Save a collection to share your selection of sources.