Search bioRxiv⌕ Search

Biology subjects

Euko, J. P.

Publications and source records attributed to Euko, J. P..

2 recordsLinked to original sources

Chemical Descriptors and Deep Learning Embeddings for Scoring de novo Peptide Designs

Peptides occupy a valuable niche between small molecules and biologics, but the clinical translation of de novo peptide designs requires rigorous scoring to simultaneously optimise target binding affinity alongside multiple developability traits, including stability, membrane permeability, aggregation propensity, and non-fouling behaviour. Here, we evaluate two distinct approaches for scoring these candidates: classical chemical descriptors and modern deep learning representations derived from protein language and folding models. Assembling nine public datasets spanning five developability traits and four binding-affinity endpoints, we find sequence-derived chemical descriptors alone contain sufficient information to predict developability task labels effectively. Given their drastically lower computational cost and higher interpretability, classical machine learning models trained on these simple descriptors frequently match or approach the performance of complex deep learning architectures, emerging as a highly efficient and interpretable alternative for high-throughput scoring. Finally, for scoring binding affinity, we demonstrate that Boltz-2 pair representations capture the most information among the tested representations; however, the model's predictive power is confounded by a significant bias from the molecular weight of the peptides. Together, these results establish a comprehensive assessment of state-of-the-art methods for predicting both peptide developability and binding affinity, highlighting the enduring value of interpretable chemical descriptors alongside deep learning in the scoring and selection of de novo peptide designs.

bioinformatics↗

CoV-UniBind: A Unified Antibody Binding Database for SARS-CoV-2

Since the emergence of SARS-CoV-2, numerous studies have investigated antibody interactions with viral variants in vitro, and several datasets have been curated to compile available protein structures and experimental measurements. However, existing data remain fragmented, limiting their utility for the development and validation of machine learning models for antibody-antigen interaction prediction. Here, we present CoV-UniBind, a unified database comprising over 75,000 entries of SARS-CoV-2 antibody-antigen sequence, binding, and structural data, integrated and standardised from three public sources and multiple peer-reviewed publications. To demonstrate its utility, we benchmarked multiple protein folding and inverse folding models across tasks relevant to antibody design and vaccine development. We expect CoV-UniBind to facilitate future computational efforts in antibody and vaccine development against SARS-CoV-2. Availability and implementationThe curated datasets, structures, model scores and antibody synonyms are free to download at https://huggingface.co/datasets/InstaDeepAI/cov-unibind. Folded structures are available upon request.

immunology↗