Search bioRxiv⌕ Search

Biology subjects

Shang, T.

Publications and source records attributed to Shang, T..

7 recordsLinked to original sources

MeDCycFold: A Rosetta Distillation Model to Accelerate Structure Prediction of Cyclic Peptides with Backbone N-methylation and D-amino Acids

Cyclic peptides with backbone N-methylated amino acids(BNMeAAs) and D-amino acids(D-AAs) have gained attention for their stability, membrane permeability, and other therapeutic potentials. Currently, Rosetta can predict their structures using energy calculations, but this method is heavily time-consuming. Moreover, structural data for cyclic peptides containing BNMeAAs and D-AAs are extremely insufficient to build a data-driven structure prediction model. To address these problems, we propose MeDCycFold, a deep learning-based Rosetta distillation model by fine-tuning the AlphaFold model. First, a cyclic peptide structure dataset is constructed using Rosetta by sampling massive conformations for cyclic peptides with BNMeAAs and D-AAs and evaluating their energy scores. Then, the AlphaFold model is fine-tuned with the extended 56 BNMeAAs and D-AAs. Besides, a relative position cyclic matrix is introduced for head-to-tail cyclization in the cyclic peptides. Finally, a force field is employed to reduce clashes in the predicted structures. Empirical experiments show that our proposed MeDCycFold speeds up structure prediction by 49 times while maintaining the prediction accuracy comparable to Rosetta, which can greatly accelerate the development of cyclic peptide drugs.

bioinformatics↗

HighPlay:Cyclic Peptide Sequence Design Based on Reinforcement Learning and Protein Structure Prediction

The structural diversity and good biocompatibility of cyclic peptides has led to their emergence as potential therapeutic agents. Traditional cyclic peptide design relies on natural template modification and combinatorial chemical library screening, but suffers from bottlenecks such as limited molecular diversity, high cost, and time-consuming.AI technologies have improved design efficiency by predicting target binding patterns and generating scaffold structures, but they still require extensive laboratory screening. In this study, we propose HighPlay, which integrates reinforcement learning (Monte Carlo Tree Search) with the HighFold structure prediction model to design cyclic peptide sequences for protein targets, dynamically exploring the sequence space without the need of predefined target information. The model was applied to the design of cyclic peptide sequences for three different targets, which were screened and verified by molecular dynamics simulation, and showed good binding affinity. Specifically, the cyclic peptide sequences designed for TEAD4 target showed micromolar-level affinity in further experimental validation.

bioinformatics↗

Predicting the structures of cyclic peptides containing unnatural amino acids by HighFold2

Cyclic peptides containing unnatural amino acids possess many excellent properties and have become promising candidates in drug discovery. Therefore, accurately predicting the three-dimensional structures of cyclic peptides containing unnatural residues will significantly advance the development of cyclic peptide-based therapeutics. Although deep learning-based structural prediction models have made tremendous progress, these models still cannot predict the structures of cyclic peptides containing unnatural amino acids. To address this gap, we introduce a novel model, HighFold2, built upon the AlphaFold-Multimer framework. HighFold2 first extends the pre-defined rigid groups and their initial atomic coordinates from natural amino acids to unnatural amino acids, thus enabling structural prediction for these residues. Then, it incorporates an additional neural network to characterize the atom-level features of peptides, allowing for multi-scale modeling of peptide molecules while enabling the distinction between various unnatural amino acids. Besides, HighFold2 constructs a relative position encoding matrix for cyclic peptides based on different cyclization constraints. Except for training using spatial structures with unnatural amino acids, HighFold2 also parameterizes the unnatural amino acids to relax the predicted structure by energy minimization for clash elimination. Extensive empirical experiments demonstrate that HighFold2 can accurately predict the three-dimensional structures of cyclic peptide monomers containing unnatural amino acids and their complexes with proteins, with the median RMSD for C reaching 1.891 [A]. All these results indicate the effectiveness of HighFold2, representing a significant advancement in cyclic peptide-based drug discovery.

bioinformatics↗

NCPepFold: Accurate Prediction of Non-canonical Cyclic Peptide Structures via Cyclization Optimization with Multigranular Representation

Artificial intelligence-based peptide structure prediction methods have revolutionized biomolecular science. However, restricting predictions to peptides composed solely of 20 natural amino acids significantly limits their practical application, as such peptides often demonstrate poor stability under physiological conditions. Here, we present NCPepFold, a computational approach that can utilize a specific cyclic position matrix to directly predict the structure of cyclic peptides with non-canonical amino acids. By integrating multi-granularity information at the residue- and atomic-level, along with fine-tuning techniques, NCPepFold significantly improves prediction accuracy, with the average peptide RMSD for cyclic peptides being 1.640 [A]. In summary, this is a novel deep learning model designed specifically for cyclic peptides with non-canonical amino acids without length restrictions, offering great potential for peptide drug design and advancing biomedical research.

bioinformatics↗

Reliable amplification of highly repetitive or low complexity sequence DNA enabled by superhelicase-mediated isothermal amplification

PCR is a cornerstone of molecular biology, but many biologically important DNA templates remain difficult to amplify. Long tandem repeats, low-complexity tracts, and sequences with extreme base composition often yield low product levels, smeared bands, stutter products, or truncated amplicons. These failures can arise because repeated cycles of thermal denaturation and reannealing promote off-register annealing, polymerase slippage, secondary-structure formation, and incomplete extension. Previously, we developed SSB-Helicase Assisted Rapid PCR (SHARP), an isothermal amplification method in which an engineered superhelicase and single-stranded DNA-binding protein replace the thermal melting step of PCR with enzymatic strand separation. Here, we tested whether SHARP can improve amplification of templates that are refractory to conventional PCR. SHARP robustly amplified up to six identical tandem repeats of the Widom 601 nucleosome-positioning sequence and up to 35 identical ankyrin repeats, targets that were poorly amplified by conventional PCR under the conditions tested. SHARP also amplified templates with extreme base composition, including a 95% AT-rich template and GC-rich templates (up to ~95% GC), as well as a 99-repeat CGG tract associated with Fragile X syndrome and a CAG/CAA repeat tract from mutant HTT exon 1 associated with Huntington disease. Together, these results show that helicase-driven isothermal amplification can expand access to repetitive and compositionally extreme DNA sequences that are challenging for conventional thermocycling-based PCR.

bioengineering↗

GAPS: Geometric Attention-based Networks for Peptide Binding Sites Identification by the Transfer Learning Approach

The identification of protein-peptide binding sites significantly advances our understanding of their interaction. Recent advancements in deep learning have profoundly transformed the prediction of protein-peptide binding sites. In this work, we describe the Geometric Attention-based networks for Peptide binding Sites identification (GAPS). The GAPS constructs atom representations using geometric feature engineering and employs various attention mechanisms to update pertinent biological features. In addition, the transfer learning strategy is implemented for leveraging the pre-trained protein-protein binding sites information to enhance training of the protein-peptide binding sites recognition, taking into account the similarity of proteins and peptides. Consequently, GAPS demonstrates state-of-the-art (SOTA) performance in this task. Our model also exhibits exceptional performance across several expanded experiments including predicting the apo protein-peptide, the protein-cyclic peptide, and the predicted protein-peptide binding sites. Overall, the GAPS is a powerful, versatile, stable method suitable for diverse binding site predictions.

bioinformatics↗

Highfold: accurately predicting cyclic peptide monomers and complexes with AlphaFold

In recent years, cyclic peptides have gained growing traction as a therapeutic modality owing to their diverse biological activities. Understanding the structures of these cyclic peptides and their complexes can provide valuable insights. However, experimental observation needs much time and money, and there still are many limitations to CADD methods. As for DL-based models, the scarcity of training data poses a formidable challenge in predicting cyclic peptides and their complexes. In this work, we present "High-fold," an AlphaFold-based algorithm that addresses this issue. By incorporating pertinent information about head-to-tailed circular and disulfide bridge structures, Highfold reaches the best performance in comparison to other various approaches. This model enables accurate prediction of cyclic peptides and their complexes, making a step to-wards resolving its structure-activity research.

bioinformatics↗