bioRxiv · 10.1101/2024.06.04.596712
Protein Sequence Domain Annotation using Language Models
Abstract
Protein domain annotation underlies large-scale functional inference and is commonly performed by scanning sequences against libraries of profile hidden Markov models (profile HMMs). We describe PSALM, a protein domain annotation method that combines (i) a pretrained protein language model (ESM-2) with (ii) a per-residue domain-state classifier and (iii) a structured probabilistic decoder that produces a single, non-overlapping set of domain calls with explicit boundaries and scores. On a benchmark of 89M protein sequences with 107M annotated domains, PSALM attains a domain-detection sensitivity-specificity tradeoff comparable to HMMER. We characterize sequence and residue-level coverage on UniProtKB, observing higher coverage for HMMER at stringent expected false positive counts (E-values) and higher coverage for PSALM at relaxed E-values. We release code for data processing, training, and inference, along with the model weights and datasets used for training, validation, and benchmarking.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sarkar, A., Krishnan, K., Eddy, S. R.. 2024-06-05. Protein Sequence Domain Annotation using Language Models. https://doi.org/10.1101/2024.06.04.596712
Cite the original work for its findings. Save a collection to share your selection of sources.