Search bioRxiv⌕ Search

Biology subjects

Abrudan, A.

Publications and source records attributed to Abrudan, A..

2 recordsLinked to original sources

Chemical Descriptors and Deep Learning Embeddings for Scoring de novo Peptide Designs

Peptides occupy a valuable niche between small molecules and biologics, but the clinical translation of de novo peptide designs requires rigorous scoring to simultaneously optimise target binding affinity alongside multiple developability traits, including stability, membrane permeability, aggregation propensity, and non-fouling behaviour. Here, we evaluate two distinct approaches for scoring these candidates: classical chemical descriptors and modern deep learning representations derived from protein language and folding models. Assembling nine public datasets spanning five developability traits and four binding-affinity endpoints, we find sequence-derived chemical descriptors alone contain sufficient information to predict developability task labels effectively. Given their drastically lower computational cost and higher interpretability, classical machine learning models trained on these simple descriptors frequently match or approach the performance of complex deep learning architectures, emerging as a highly efficient and interpretable alternative for high-throughput scoring. Finally, for scoring binding affinity, we demonstrate that Boltz-2 pair representations capture the most information among the tested representations; however, the model's predictive power is confounded by a significant bias from the molecular weight of the peptides. Together, these results establish a comprehensive assessment of state-of-the-art methods for predicting both peptide developability and binding affinity, highlighting the enduring value of interpretable chemical descriptors alongside deep learning in the scoring and selection of de novo peptide designs.

bioinformatics↗

Model Agnostic Conditioning of Boltzmann Generators for Peptide Cyclization

Macrocyclic peptides offer strong therapeutic potential due to their enhanced binding affinity and protease resistance, but their design remains a challenge due to limited structural data and tools that address only a narrow set of cyclization chemistries. Moreover, existing models are built to only consider ground state or mean conformations, rather than conformational ensembles that more accurately describes peptides. We introduce CO_SCPLOWYCC_SCPLOWLOPS (a Cyclic Loss for the Optimization of Peptide Structures), a model-agnostic framework that conditions Boltzmann generators to sample valid cyclic conformations--without retraining. To overcome the scarcity of cyclic peptide data, we reformulate the design problem in terms of conditional sampling over linear peptide structures via chemically informed loss functions. CO_SCPLOWYCC_SCPLOWLOPS encompasses 18 possible inter-amino acid crosslinks enabled by 6 diverse chemical reactions, and is readily extensible to many more. It leverages tetrahedral geometry constraints, using six interatomic distances to define a kernel density-estimated joint distribution from MD simulations. We demonstrate CO_SCPLOWYCC_SCPLOWLOPSs versatility via two distinct generative models: a modified Sequential Boltzmann Generator (SBG) (Tan et al., 2025) and the Equivariant Normalizing flow (ECNF) of Klein & Noe (2024). In both settings, CO_SCPLOWYCC_SCPLOWLOPS successfully biases the Boltzmann distribution toward chemically plausible macrocycles.

biophysics↗