Search bioRxiv⌕ Search

Biology subjects

Chronowska, M.

Publications and source records attributed to Chronowska, M..

2 recordsLinked to original sources

MHChron: diversity-balanced dataset design for robust peptide-MHC binding prediction across MHC class I and II

Accurate prediction of peptide-MHC (pMHC) binding is central to immunogenicity assessment, yet many existing predictors are trained and evaluated on narrow allele sets and restricted peptide-lengths. Here, we present MHChron, a unified pMHC binding prediction framework predicated on systematic data curation, meticulous engineering of dataset balance and diversity, and rigorous evaluation through careful splits controlling for data leakage. We assemble one of the most diverse pMHC training dataset reported to date, integrating publicly available binding data across a broad allele coverage (class I n=214, class II n=98) and peptide length range (from 8 to 36 residues). Using a focused and carefully sampled subset of this dataset, we train complementary sequence-based and structure-aware models and test them under increasingly stringent generalisation regimes. Both models achieve consistently strong performance, outperforming the evaluated state-of-the-art predictors despite being trained on numerically fewer data points. Notably, the structure-aware model did not consistently surpass the sequence-based model, except under the most demanding setting of extrapolation to unseen allele clusters, suggesting that performance gains stem primarily from dataset diversity and rigorous evaluation rather than architectural complexity. Sequence-based MHChron is released with reproducible installation and an automated whole-protein screening pipeline, enabling broad and practical use.

bioinformatics↗

The Protein Design Archive (PDA): insights from 40 years of protein design

The field of protein design has changed dramatically over the last 40 years, with a range of methods developing from rational design to more recent data-driven approaches. While considerable insight could be gained from analysing designed proteins, there is no single resource that brings together all the relevant data, making it difficult to identify current challenges and opportunities within the field. Here we present the Protein Design Archive, a website and database of designed proteins. Using the database, we performed systematic analysis that reveals a rapid increase in the number and complexity of designs over time and uncovers biases in their amino-acid usage and secondary structure content. Finally, sequence- and structure-based analysis demonstrates the breadth and novelty of designs within the archive. The PDA will be a valuable resource for guiding development of protein design, helping the field reach its potential. The PDA is available freely and without registration at https://pragmaticproteindesign.bio.ed.ac.uk/pda/.

bioinformatics↗