Search bioRxiv⌕ Search

Biology subjects

Ouellet, S.

Publications and source records attributed to Ouellet, S..

2 recordsLinked to original sources

DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active state complexes and deep learning embeddings

Peptide-activated G protein-coupled receptors (GPCRs) regulate critical physiological processes such as metabolism, neural signalling, and endocrine function through their interaction with neuropeptides and peptide hormones. Despite their importance, identifying endogenous peptide agonists for GPCRs remains challenging, particularly for orphan receptors without known ligands. Recent advances in deep learning-based protein structure prediction, exemplified by AlphaFold (AF), have shown application beyond structural modelling, including protein-protein interaction prediction. Given that GPCR-peptide agonist interactions represent a specialized form of protein-protein interaction, we leveraged a dataset of experimentally validated agonist and non-agonist GPCR-peptide interactions from Caenorhabditis elegans to evaluate AF-Multimers ability to distinguish agonist-bound complexes. When modelling GPCR-peptide complexes, AF-Multimer confidence metrics partially discriminate agonist from non-agonist interactions, with improved discrimination achieved by utilizing AF-Multistate-derived active-state templates. We further investigated whether embeddings from the hidden layer of AF-Multimers neural network could distinguish agonist from non-agonist complexes. Feature performance analysis reveals that AF- Multimers pair representations outperform single representations, with distinct subregions of the pair representation providing complementary predictive signals. Building on these insights, we developed DeorphaNN, a graph neural network integrating active-state GPCR-peptide structural predictions, interatomic interactions, and deep learning embeddings to prioritize GPCR-peptide agonist interactions. DeorphaNN generalizes across datasets derived from different species, including annelids and humans, and successfully uncovered peptide agonists for two orphan GPCRs. DeorphaNN offers a novel computational resource to accelerate deorphanization by prioritizing GPCR-peptide agonist candidates for AI-guided experimental validation.

bioinformatics↗

CysPresso: Prediction of cysteine-dense peptide expression in mammalian cells using deep learning protein representations

Background: Cysteine-dense peptides (CDPs) are an attractive pharmaceutical scaffold that display extreme biochemical properties, low immunogenicity, and the ability to bind targets with high affinity and selectivity. While many CDPs have potential and confirmed therapeutic uses, synthesis of CDPs is a challenge. Recent advances have made the recombinant expression of CDPs a viable alternative to chemical synthesis. Moreover, identifying CDPs that can be expressed in mammalian cells is crucial in predicting their compatibility with gene therapy and mRNA therapy. Currently, we lack the ability to identify CDPs that will express recombinantly in mammalian cells without labour intensive experimentation. To address this, we developed CysPresso, a novel machine learning model that predicts recombinant expression of CDPs based on primary sequence. Results: We tested various protein representations generated by deep learning algorithms (SeqVec, proteInfer, AlphaFold2) for their suitability in predicting CDP expression and found that AlphaFold2 representations possessed the best predictive features. We then optimized the model by concatenation of AlphaFold2 representations, time series transformation with random convolutional kernels, and dataset partitioning. Conclusion: Our novel model, CysPresso, is the first to successfully predict recombinant CDP expression in mammalian cells and is particularly well suited for predicting recombinant expression of knottin peptides. When preprocessing the deep learning protein representation for supervised machine learning, we found that random convolutional kernel transformation preserves more pertinent information relevant for predicting expressibility than embedding averaging. Our study showcases the applicability of deep learning-based protein representations, such as those provided by AlphaFold2, in tasks beyond structure prediction.

bioinformatics↗