Search bioRxiv⌕ Search

Biology subjects

McWhite, C.

Publications and source records attributed to McWhite, C..

1 recordsLinked to original sources

Cross-Attention Over RNA And Protein Sequences Enables Generalizable Interaction Prediction

Computational predictions are essential to characterize the RNA-protein interaction landscape, yet a persistent gap between benchmark performance and practical utility suggests that current models have limited generalization capabilities. To address this issue, we present CORAL (Cross-attention for RNA-protein Association Learning), a deep learning framework for the prediction of RNA-protein interactions that integrates pretrained protein (ESM-2) and RNA (DNABERT2) language models through bidirectional cross-attention with Low-Rank Adaptation fine-tuning. We also introduce a benchmarking framework that rigorously addresses the problem of data redundancy between training and test sets, which greatly inflates model performances reported in the literature. To this end we adopt three partitioning strategies of increasing stringency: conventional random splits, pairwise non-redundant splits, and component-wise non-redundant splits. CORAL maintains an F1 score of 0.65 under the most stringent component-wise evaluation, compared to 0.47 for the next-best method retaining discriminative behavior. Interpretability analyses further reveal that the cross-attention mechanism captures biologically meaningful features of molecular recognition at two complementary structural scales: at atomic resolution, specific attention heads systematically attend to structurally defined contact positions, showing 26% elevated attention at interface residues across 309 experimentally resolved complexes (p < 0.001); at the domain level, protein-side attention localizes to annotated RNA-binding domains across 94 % of 462 proteins examined (median within-protein Cohens d {approx} 0.94). Together, these findings establish that current RPI prediction benchmarks substantially inflate performance estimates and demonstrate that cross-modal attention architectures yield improved generalization alongside mechanistically interpretable representations.

bioinformatics↗