Search bioRxiv⌕ Search

Biology subjects

Deibler, K.

Publications and source records attributed to Deibler, K..

3 recordsLinked to original sources

PeptideMTR: Scaling SMILES-Based Language Models for Therapeutic Peptide Engineering

Therapeutic peptides occupy a unique middle ground in drug discovery, offering the high specificity of protein interactions with the chemical diversity of small molecules, yet they currently fall in a computational blind spot. Existing foundation models cannot handle them effectively: protein models are restricted to natural amino acids, while chemical models struggle to process large, polymer-like sequences. This disconnect has forced the field to rely on static chemical descriptors that fail to capture subtle chemical details or on complex multi-embedding pipelines that are custom tailored to specific datasets. To bridge this gap, we present PeptideCLM-2, a suite of chemical language models trained on over 100 million molecules to natively represent complex peptide chemistry. This modeling approach expands the available toolkit of machine learning models for therapeutic peptides. Benchmarking results show strong performance versus prior methods for predicting development endpoints including membrane diffusion, biological function, and half life.

bioinformatics↗

Predicting peptide aggregation with protein language model embeddings

Amyloid fibrils, a form of peptide aggregate, are associated with multiple diseases and hinder the development of therapeutics. The experimental characterization of aggregating peptides is resource-intensive and data are scarce, limiting the development of accurate models. We present a deep-learning model, PALM (Predicting Aggregation with Language Model embeddings), which uses transfer learning to predict aggregation from embeddings extracted from a pretrained protein language model (pLM). PALM is trained on the WaltzDB-2.0 dataset to classify peptides and identify aggregation-prone regions within a sequence at single-residue resolution. Compared to existing models, it exhibits competitive performance on diverse held-out experimental datasets. We find that PALM fails to identify single mutations that increase the rate of aggregation of amyloid beta peptide; however, training the PALM architecture on a larger dataset, CANYA NNK1-3, substantially improves performance in this task. These results show that transfer learning with pLM embeddings improves performance when training on small datasets, but highlight that challenging tasks, such as predicting the effect of single mutations, require more experimental data.

bioinformatics↗

De novo design of miniprotein agonists and antagonists targeting G protein-coupled receptors

G protein-coupled receptors (GPCRs) play key roles in physiology and are central targets for drug discovery and development, but the design of protein agonists and antagonists has been challenging as GPCRs are integral membrane proteins and conformationally dynamic. Here we describe computational de novo design methods and a high throughput "receptor diversion" microscopy-based screen for generating GPCR binding miniproteins with high affinity, potency and selectivity, and the use of these methods to generate agonists for MRGPRX1, NK1R and CCR5, as well as antagonists for CXCR4, CCR5, OXTR, GLP1R, GIPR, GCGR, PTH1R and CGRPR.. Cryo-electron microscopy data reveals atomic-level agreement between designed and experimentally determined structures for CGRPR- and CXCR4-bound antagonists and MRGPRX1-bound agonists. Our de novo design and screening approach opens new frontiers in GPCR drug discovery and development.

bioengineering↗