Search bioRxiv⌕ Search

Biology subjects

Motmaen, A.

Publications and source records attributed to Motmaen, A..

5 recordsLinked to original sources

De novo design of phospho-tyrosine peptide binders

Phosphorylation on tyrosine is a key step in many signaling pathways. Despite recent progress in de novo design of protein binders, there are no current methods for designing binders that recognize phosphorylated proteins and peptides; this is a challenging problem as phosphate groups are highly charged, and phosphorylation often occurs within unstructured regions. Here we introduce RoseTTAFold Diffusion 2 for Molecular Interfaces (RFD2-MI), a deep generative framework for the design of binders for protein, ligand, and covalently modified protein targets. We demonstrate the power and versatility of this method by designing binders for four critical phosphotyrosine sites on three clinically relevant targets: Cluster of Differentiation 3 (CD3{varepsilon}), Epidermal Growth Factor Receptor (EGFR), Insulin Receptor (INSR) and Signal Transducer and Activator of Transcription 5 (STAT5). Experimental characterization shows that the designs bind their phosphotyrosine containing targets with affinities comparable to native binding sites and have negligible binding to non-phosphorylated targets or phosphopeptides with different sequences. X-ray crystal structures of generated binders to CD3{varepsilon} and EGFR are very close to the design models, demonstrating the accuracy of the design approach. A designed binder to an EGFR intracellular region phosphorylated upon EGF activation co-localizes with the receptor following EGF stimulation in single-particle tracking (SPT) experiments, demonstrating pY specific recognition in living cells. RFD2-MI provides a generalizable all-atom diffusion framework for probing and modulating phosphorylation-dependent signaling, and more generally, for developing research tools and targeted therapeutics against post-translationally modified proteins.

bioengineering↗

Accelerating protein design by scaling experimental characterization

Recent advances in de novo protein design have greatly outpaced standard protein biochemistry workflows, and experimental testing has been a bottleneck in the validation of new designs and methodologies. Here, we describe experimental and computational workflows to address the issues of scale, speed and reproducibility of common in vitro protein testing methods, enabling at least an order of magnitude increase in throughput while reducing wetlab time. Semi-Automated Protein Production (SAPP) is a rapid, modular, scalable and cost-effective protocol, enabling up to milligram-scale protein production, and standardized characterization including yield, dispersity, and oligomeric state of hundreds of designs per day, at the cost-equivalent of a few DNA oligos per construct. End-to-end protocol execution takes 48 hours, with about 6 hours spent benchside using mostly standard laboratory equipment. This protocol has become the standard at our institute, providing critical experimental validation for dozens of projects spanning tens of thousands of designs. We showcase the power of the platform by using it to rapidly characterize de novo designed inhibitors of respiratory syncytial virus. Since at least 80% of SAPPs total cost comes from synthetic DNA, we also developed a scalable demultiplexing protocol (DMX) to leverage oligo pools as input DNA, providing a further 5-fold reduction in costs, enabling >1000 designs to be purified and characterized in arrayed, clonal format at a cost of $5 per construct. By reframing standard molecular biology practices and orchestrating wetlab workflows with partial automation instead of complex end-to-end robotics, these protocols should be widely adoptable, accelerating protein design.

biochemistry↗

De novo design of miniprotein agonists and antagonists targeting G protein-coupled receptors

G protein-coupled receptors (GPCRs) play key roles in physiology and are central targets for drug discovery and development, but the design of protein agonists and antagonists has been challenging as GPCRs are integral membrane proteins and conformationally dynamic. Here we describe computational de novo design methods and a high throughput "receptor diversion" microscopy-based screen for generating GPCR binding miniproteins with high affinity, potency and selectivity, and the use of these methods to generate agonists for MRGPRX1, NK1R and CCR5, as well as antagonists for CXCR4, CCR5, OXTR, GLP1R, GIPR, GCGR, PTH1R and CGRPR.. Cryo-electron microscopy data reveals atomic-level agreement between designed and experimentally determined structures for CGRPR- and CXCR4-bound antagonists and MRGPRX1-bound agonists. Our de novo design and screening approach opens new frontiers in GPCR drug discovery and development.

bioengineering↗

Design of high specificity binders for peptide-MHC-I complexes

Class I MHC molecules present peptides derived from intracellular antigens on the cell surface for immune surveillance, and specific targeting of these peptide-MHC (pMHC) complexes could have considerable utility for treating diseases. Such targeting is challenging as it requires readout of the few outward facing peptide antigen residues and the avoidance of extensive contacts with the MHC carrier which is present on almost all cells. Here we describe the use of deep learning-based protein design tools to de novo design small proteins that arc above the peptide binding groove of pMHC complexes and make extensive contacts with the peptide. We identify specific binders for ten target pMHCs which when displayed on yeast bind the on-target pMHC tetramer but not closely related peptides. For five targets, incorporation of designs into chimeric antigen receptors leads to T-cell activation by the cognate pMHC complexes well above the background from complexes with peptides derived from proteome. Our approach can generate high specificity binders starting from either experimental or predicted structures of the target pMHC complexes, and should be widely useful for both protein and cell based pMHC targeting.

immunology↗

Peptide binding specificity prediction using fine-tuned protein structure prediction networks

Peptide binding proteins play key roles in biology, and predicting their binding specificity is a long-standing challenge. While considerable protein structural information is available, the most successful current methods use sequence information alone, in part because it has been a challenge to model the subtle structural changes accompanying sequence substitutions. Protein structure prediction networks such as AlphaFold model sequence-structure relationships very accurately, and we reasoned that if it were possible to specifically train such networks on binding data, more generalizable models could be created. We show that placing a classifier on top of the AlphaFold network and fine-tuning the combined network parameters for both classification and structure prediction accuracy leads to a model with strong generalizable performance on a wide range of Class I and Class II peptide-MHC interactions that approaches the overall performance of the state-of-the-art NetMHCpan sequence-based method. The peptide-MHC optimized model shows excellent performance in distinguishing binding and non-binding peptides to SH3 and PDZ domains. This ability to generalize well beyond the training set far exceeds that of sequence only models, and should be particularly powerful for systems where less experimental data is available. Significance statementPeptide binding proteins carry out a variety of biological functions in cells and predicting their binding specificity could significantly improve our understanding of molecular pathways. Deep neural networks have achieved high structure prediction accuracy, but are not trained to predict binding specificity. Here we describe an approach to extending such networks to jointly predict protein structure and binding specificity. We incorporate AlphaFold into this approach, and fine-tune its parameters on peptide-MHC Class I and II structural and binding data. The fine-tuned model approaches state-of-the-art classification accuracy on peptide-MHC specificity prediction and generalizes to other peptide-binding systems such as the PDZ and SH3 domains.

bioinformatics↗