Search bioRxiv⌕ Search

Biology subjects

Zhen, Q.

Publications and source records attributed to Zhen, Q..

2 recordsLinked to original sources

Highly Accurate Estimation of the Fold Accuracy of Protein Structural Models

BackgroundThe function of a protein is intrinsically linked to its three-dimensional fold, and deep learning has revolutionized the field by enabling high-accuracy structure prediction at an unprecedented scale. Nevertheless, the growing deployment of these predictive pipelines in drug discovery and structural biology reveals a critical bottleneck that lies in the lack of independent and rigorous estimation of model accuracy (EMA) methodologies. ResultsHere we present DeepUMQA-Global, a single-model deep learning framework for estimating accuracy of protein structure models. Our method employs a structure-sequence cross-consistency mechanism to evaluate the bidirectional compatibility between the predicted structure and the input sequence, enabling comprehensive characterization of fold accuracy. DeepUMQA-Global outperforms the self-assessment confidence scores of AlphaFold3, achieving improvements of 57.8% in Pearson correlation and 49.0% in Spearman correlation. With respect to the CASP16 retrospective benchmark, DeepUMQA-Global outperforms all single-model accuracy estimation methods that participated in CASP16 and achieves performance comparable to that of the top consensusObased methods. A lightweight consensus strategy built upon DeepUMQA-Global ranks first among all CASP16 participants, surpassing all other methods, including consensus approaches, and highlighting the strength of our method. Remarkably, DeepUMQA-Global demonstrates a strong ability to discriminate between alternative conformational states of proteins, as evidenced in the CASP unique alternative conformation protein complex target and the CoDNaS benchmark. ConclusionsOur results indicate that DeepUMQA-Global can be extended to broader protein modeling tasks, moving beyond static evaluation to offer a foundation for dynamic conformation EMA, where it accurately discriminates alternative conformational states and demonstrates reliable predictive fidelity in model accuracy estimation.

bioinformatics↗

ProSiteHunter: A unified framework for sequence-based prediction of protein-nucleic acid and protein-protein binding sites

Accurate prediction of protein binding sites is essential for elucidating protein function, understanding molecular interaction mechanisms, and facilitating drug design. However, existing sequence-based approaches are often designed for specific binding-site types and therefore lack generality, whereas structure-based methods typically rely on high-quality structural models, limiting their applicability. Here, we introduce ProSiteHunter, a unified sequence-based framework for protein binding-site prediction, which integrates a fine-tuned protein language model (SiteT5) with a multi-source feature-fusion network that incorporates evolutionary, geometric, and statistical features, while employing bidirectional semantics, local associations, and global dependencies for comprehensive binding-site characterization. The method was systematically evaluated on diverse binding sites prediction tasks, where ProSiteHunter achieved a 39.1% average improvement in PRAUC for protein-DNA/RNA/protein tasks and a 7.4% PRAUC enhancement on the particularly challenging antibody-antigen task over state-of-the-art methods. Moreover, ProSiteHunter is capable of identifying local flexible sites that complement AlphaFold3 predictions and improving the accuracy of antibody-antigen interaction prediction. These results highlight ProSiteHunter as an efficient and unified approach for accurate and robust prediction of diverse protein binding sites.

bioinformatics↗