Search bioRxiv⌕ Search

Biology subjects

Trolliet, Q.

Publications and source records attributed to Trolliet, Q..

2 recordsLinked to original sources

Chemical Descriptors and Deep Learning Embeddings for Scoring de novo Peptide Designs

Peptides occupy a valuable niche between small molecules and biologics, but the clinical translation of de novo peptide designs requires rigorous scoring to simultaneously optimise target binding affinity alongside multiple developability traits, including stability, membrane permeability, aggregation propensity, and non-fouling behaviour. Here, we evaluate two distinct approaches for scoring these candidates: classical chemical descriptors and modern deep learning representations derived from protein language and folding models. Assembling nine public datasets spanning five developability traits and four binding-affinity endpoints, we find sequence-derived chemical descriptors alone contain sufficient information to predict developability task labels effectively. Given their drastically lower computational cost and higher interpretability, classical machine learning models trained on these simple descriptors frequently match or approach the performance of complex deep learning architectures, emerging as a highly efficient and interpretable alternative for high-throughput scoring. Finally, for scoring binding affinity, we demonstrate that Boltz-2 pair representations capture the most information among the tested representations; however, the model's predictive power is confounded by a significant bias from the molecular weight of the peptides. Together, these results establish a comprehensive assessment of state-of-the-art methods for predicting both peptide developability and binding affinity, highlighting the enduring value of interpretable chemical descriptors alongside deep learning in the scoring and selection of de novo peptide designs.

bioinformatics↗

Mind the Gap: An Embedding Guide to Safely Travel in Sequence Space

We present a hybrid approach combining a protein language model (pLM) with Monte Carlo (MC) sampling for generating enzyme mutants free of mutations deleterious for structural preservation. Given the amino acid sequence of the original enzyme and a set of residues for which the local environment should be conserved, i.e., the catalytic site, our approach generates mutants that differ vastly in the overall sequence while retaining the geometry of the conserved region, thereby representing promising candidates for further experimental screening. Unlike end-to-end deep learning approaches, whose results are harder to interpret and control, the use of a well-established, classic technique such as MC sampling allows us to easily interpret the generative process as the sampling of an energy landscape determined by the pLM. In turn, such an interpretation enables us to steer this generative process and control its outcome by making use of robust statistical mechanics concepts, e.g., temperature, thereby explicitly guaranteeing certain properties of the generated mutants. Given the increasing relevance of generative algorithms in the design and search for novel, optimised enzymes, we believe that our results constitute an important step for the future development of this class of techniques. To facilitate experimental verification, we finally provide hundreds of sequences for 13 different enzymes involved in catalytic processes ranging from carbon dioxide conversion to DNA replication.

synthetic biology↗