k-mer Distributions of Aminoacid Sequences are Optimised Across the Proteome
k-mer based methods are widely utilized for the analysis of nucleotide sequences and were successfully applied to proteins in several works. However, the reasons for the species-specificity of aminoacid k-mer distributions are unknown. In this work I show that performance of these methods is not only due to orthology between k-mers in different proteomes, which implies the existence of some factors optimizing k-mer distributions of proteins in a species-specific manner. Whatever these factors could be, they are affecting most if not all proteins and are more pronounced in structurally organized regions.
bioinformatics↗