bioRxiv · 10.64898/2026.02.02.703349
EvoPool: Evolution-Guided Pooling of Protein Language Model Embeddings
Abstract
Protein language models (PLMs) encode amino acid sequences into residue-level embeddings that must be pooled into fixed-size representations for downstream protein-level prediction tasks. Although these embeddings implicitly reflect evolutionary constraints, existing pooling strategies operate on single sequences and do not explicitly leverage information from homologous sequences or multiple sequence alignments. We introduce EvoPool, a self-supervised pooling framework that integrates evolutionary information from homologs directly into aggregated PLM representations using optimal transport. Our method constructs a fixed-size evolutionary anchor from an arbitrary number of homologous sequences and uses sliced Wasserstein distances to derive query protein embeddings that are geometrically informed by homologous sequence embeddings. Experiments across multiple state-of-the-art PLM families on the ProteinGym benchmark show that EvoPool consistently outperforms standard pooling baselines for variant effect prediction, demonstrating that explicit evolutionary guidance substantially enhances the functional utility of PLM representations. Our implementation code is available at https://github.com/navid-naderi/EvoPool.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
NaderiAlizadeh, N., Singh, R.. 2026-02-04. EvoPool: Evolution-Guided Pooling of Protein Language Model Embeddings. https://doi.org/10.64898/2026.02.02.703349
Cite the original work for its findings. Save a collection to share your selection of sources.