bioRxiv · 10.1101/2024.06.05.597668
Alignment of multiple protein sequences without using amino acid frequencies.
Abstract
Current algorithms for aligning protein sequences use substitutability scores that combine the probability to find an amino acid in a specific pair of amino acids and marginal probability to find this amino acid in any pair. However, the positional probability of finding the amino acid at a place in alignment is also conditional on the amino acids at the sequence itself. Content-dependent corrections overparameterize protein alignment models. Here, we propose an approach that is based on (dis)similarily measures, which do not use the marginal probability, and score only probabilities of finding amino acids in pairs. The dissimilarity scoring matrix endows a metric space on the set of aligned sequences. This allowed us to develop new heuristics. Our aligner does not use guide trees and treats all sequences uniformly. We suggest that such alignments that are done without explicit evolution-based modeling assumptions should be used for testing hypotheses about evolution of proteins (e.g., molecular phylogenetics).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shirokov, R., Shelyekhova, V.. 2024-06-09. Alignment of multiple protein sequences without using amino acid frequencies.. https://doi.org/10.1101/2024.06.05.597668
Cite the original work for its findings. Save a collection to share your selection of sources.