Search bioRxiv⌕ Search

Biology subjects

Pai, N.

Publications and source records attributed to Pai, N..

2 recordsLinked to original sources

Disentangling RNA evolution and thermodynamics in genomic language models

Genomic language models (gLMs) trained only on large-scale nucleic acid sequence data seem to capture signals of RNA structure, yet the specifics of how remain unclear. Using the categorical Jacobian (CJ) operation, a model-agnostic operation for querying pairwise dependencies, we systematically compared three flagship gLMs: RNA-FM, Evo 2, and gLM2. We found that CJ signals recover base pairs supported by evolutionary covariation analyses, consistent with findings in protein language models. Surprisingly, CJ also recovers base pairs lacking evolutionary support but predicted by biophysical nearest-neighbor models. Is it possible gLMs have "learned" RNA thermodynamics? We noticed nearest-neighbor RNA folding models often predict reflected structures when given reversed sequences, consistent with these models modular and grammar-like nature. We leveraged this observation to create a simple "mirror test" that we found gLMs routinely fail, indicating they have not learned generalizable biophysics-based rules for RNA structure. Nevertheless, their apparent thermodynamic signal potentially confounds interpreting gLM pairwise dependencies as evidence of evolutionary conservation. We therefore introduce a method using synthetic sequences as a control for detecting significant learned signal. Our results demonstrate that gLMs can mimic thermodynamics through learned sequence context rather than general physical principles, but solutions exist for disentangling patterns in language models.

biophysics↗

Putting computational models of immunity to the test - an invited challenge to predict B. pertussis vaccination outcomes

Systems vaccinology studies have been used to build computational models that predict individual vaccine responses and identify the factors contributing to differences in outcome. Comparing such models is challenging due to variability in study designs. To address this, we established a community resource to compare models predicting B. pertussis booster responses and generate experimental data for the explicit purpose of model evaluation. We here describe our second computational prediction challenge using this resource, where we benchmarked 49 algorithms from 53 scientists. We found that the most successful models stood out in their handling of nonlinearities, reducing large feature sets to representative subsets, and advanced data preprocessing. In contrast, we found that models adopted from literature that were developed to predict vaccine antibody responses in other settings performed poorly, reinforcing the need for purpose-built models. Overall, this demonstrates the value of purpose-generated datasets for rigorous and open model evaluations to identify features that improve the reliability and applicability of computational models in vaccine response prediction.

immunology↗