Search bioRxiv⌕ Search

Biology subjects

Polunina, P. V.

Publications and source records attributed to Polunina, P. V..

2 recordsLinked to original sources

Tree-aware conditional language modeling recovers mutational patterns of viral evolution

The evolutionary accessibility of a protein mutation depends on the sequence background in which it arises and its lineage history, yet most protein language models estimate sequence plausibility without explicitly considering the ordered sequence changes through which descendants arise. We developed evoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features. We demonstrated our approach using SARS-CoV-2 spike protein, pairing sequences from early Omicron lineages according to their positions on a mutation-annotated phylogeny, and evaluating model performance on sequence pairs from later lineages. Prompt-masking experiments showed that incorporating phylogenetic context substantially increased reliance on the supplied input information compared with a sequence-only model. Generated descendant sequences accurately reproduced the positional distribution of mutations observed during viral evolution, with strong correlations between predicted and observed mutation-frequency profiles for both the receptor-binding domain (Spearman's {rho} = 0.823) and the full spike protein ({rho} = 0.736). Although prediction accuracy for individual substitutions decreased with increasing evolutionary distance, the model consistently captured aggregate mutational patterns across the spike protein. Model-assigned mutation probabilities were also enriched among substitutions experimentally tolerated in deep mutational scanning assays of Omicron BA.2 receptor-binding domain expression (1.19-fold enrichment) and ACE2 binding (1.04-fold enrichment), despite the model being trained solely on observed ancestor-descendant sequence pairs and associated phylogenetic context features. These results demonstrate that explicitly providing protein language models with phylogenetic context during sequence generation can recover lineage-specific mutational patterns and yields probabilistic predictions consistent with experimentally measured functional constraints. evoPLM-Tree provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.

bioinformatics↗

VEFill: a model for accurate and generalizable deep mutational scanning score imputation across protein domains

BackgroundDeep Mutational Scanning (DMS) assays can systematically assess the effects of amino acid substitutions on protein function. While DMS datasets have been generated for many targets, they often suffer from incomplete variant coverage due to technical constraints, limiting their utility in variant interpretation and downstream analyses. ResultsWe developed VEFill, a gradient boosting model for imputing missing DMS scores across protein domains. VEFill is trained on the Human Domainome 1 dataset, a large, standardized set of DMS experiments using a uniform stability-based assay, and integrates a broad set of additional biologically informative features including ESM-1v sequence embeddings, evolutionary conservation (EVE scores), amino acid substitution matrices, and physicochemical descriptors. The model achieved robust predictive performance (R2 = 0.64, Pearson r = 0.80). It also demonstrated reliable generalization to unseen proteins in other stability-based datasets, while showing weaker performance on activity-based assays. Per-protein models further confirmed VEFills effectiveness under limited-data conditions. A reduced two-feature version using only ESM-1v embeddings and mean DMS scores performed comparably to the full model, suggesting a computationally efficient alternative. However, true zeroshot prediction without positional context remains a challenge, particularly for functionally complex proteins. ConclusionsVEFill offers an interpretable, scalable framework for DMS score imputation, especially effective in stability-focused and sparse-data settings. It enables systematic mutation prioritization and may support the design of efficient experimental libraries for variant effect studies.

bioinformatics↗