Search bioRxiv⌕ Search

Biology subjects

Pugh, C. W. J.

Publications and source records attributed to Pugh, C. W. J..

2 recordsLinked to original sources

Blending physics-based and inverse folding models to disentangle variant effects on stability and function

Protein sequences are constrained not only by the need to fold into stable structures, but also by specific functional requirements imposed by natural selection. Yet predictions of how amino-acid changes affect proteins typically collapse these constraints into a single scalar score. Quantitatively separating these effects at scale remains an open challenge, with direct relevance spanning protein design to understanding the molecular mechanisms of disease. Inverse-folding (IF) models have emerged as fast, unsupervised predictors of folding energy changes ({Delta}{Delta}G), but because they learn statistical correspondences between structure and sequence, they can conflate conservation driven by function with conservation driven by stability. Here, we show that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental {Delta}{Delta}G and, crucially, reduces IF model bias at functional sites. Applying the best-performing blend together with an evolutionary language model, we decompose each variants evolutionary cost into folding energy and dark energy, the latter capturing functional constraints beyond folding stability. With this decomposition, and without the need for supervision, we find that disease gain-of-function variants show a distinct functional signature from loss-of-function variants. In particular, we identify oncogenic drivers as largely preserving stability while exhibiting high dark energy, as opposed to tumor suppressors which are predominantly destabilized, paving the way to a mechanistic understanding of driver mutations in cancer. Together, these results provide a scalable framework for accurate {Delta}{Delta}G prediction and mechanistic disentanglement of variant effects.

biophysics↗

From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language Models

Generative models trained on natural sequences are increasingly used to predict the effects of genetic variation, enabling progress in therapeutic design, disease risk prediction, and synthetic biology. In the zero-shot setting, variant impact is estimated by comparing the likelihoods of sequences, under the assumption that likelihood serves as a proxy for fitness. However, this assumption often breaks down in practice: sequence likelihood reflects not only evolutionary fitness constraints, but also phylogenetic structure and sampling biases, especially as model capacity increases. We introduce Likelihood-Fitness Bridging (LFB), a simple and general strategy that improves variant effect prediction by averaging model scores across sequences subject to similar selective pressures. Assuming an Ornstein-Uhlenbeck model of evolution, LFB can be viewed as a way to marginalize the effects of genetic drift, although its benefits appear to extend more broadly. LFB applies to existing protein and genomic language models without requiring retraining, and incurs only modest computational overhead. Evaluated on large-scale deep mutational scans and clinical benchmarks, LFB consistently improves predictive performance across model families and sizes. Notably, it reverses the performance plateau observed in larger protein language models, making the largest models the most accurate when combined with LFB. These results suggest that accounting for phylogenetic and sampling biases is essential to realizing the full potential of large sequence models in variant effect prediction.

bioinformatics↗