bioRxiv · 10.1101/2025.08.06.668973
FusionProt: Fusing Sequence and Structural Information for Unified Protein Representation Learning
Abstract
Accurate protein representations that integrate sequence and three-dimensional (3D) struc-ture are critical to many biological and biomedical tasks. Most existing models either ignore structure or combine it with sequence through a single, static fusion step. Here we present FusionProt, a unified model that learns representations via iterative, bidirectional fusion be-tween a protein language model and a structure encoder. A single learnable token serves as a carrier, alternating between sequence attention and spatial message passing across layers. FusionProt is evaluated on Enzyme Commission (EC), Gene Ontology (GO), and mutation stability prediction tasks. It improves Fmax by a median of +1.3 points (up to +2.0) across EC and GO benchmarks, and boosts AUROC by +3.6 points over the strongest baseline on mutation stability. Inference cost remains practical, with only [~] 2-5% runtime over-head. Beyond state-of-the-art performance, we further demonstrate FusionProts practical relevance through representative biological case studies, suggesting that the model captures biologically relevant features.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kalifa, D., Singer, U., Radinsky, K.. 2025-08-08. FusionProt: Fusing Sequence and Structural Information for Unified Protein Representation Learning. https://doi.org/10.1101/2025.08.06.668973
Cite the original work for its findings. Save a collection to share your selection of sources.