Search bioRxiv⌕ Search

Biology subjects

Jayasinghe-Arachchige, V. M.

Publications and source records attributed to Jayasinghe-Arachchige, V. M..

2 recordsLinked to original sources

Structure-Based Classification of CRISPR/Cas9 Proteins: A Machine Learning Approach to Elucidating Cas9 Allostery

The CRISPR/Cas9 system is a powerful gene-editing tool. Its specificity and stability rely on complex allosteric regulation. Understanding these allosteric regulations is essential for developing high-fidelity Cas9 variants with reduced off-target effects. Here, we introduce a novel structure-based machine learning (ML) approach to systematically identify long-range allosteric networks in Cas9. Our ML model was trained using all available Cas9 structures, ensuring a comprehensive representation of Cas9s structural landscape. We then applied this model to Streptococcus pyogenes Cas9 (SpCas9) to demonstrate the feature selection process. Using C-C inter-residue distances, we mapped key allosteric networks and refined them through a two-stage SHAP feature selection (FS) strategy, reducing a vast feature space to 28 critical Lysine-Arginine (Lys-Arg) residue pairs that mediate SpCas9 interdomain communication, stability, and specificity. These Lys-Arg pairs initially shared a 46.5 [A] inter-residue distance, but molecular dynamics simulations revealed distinct stabilization behaviors, indicating a hierarchical allosteric network. Further mutational analysis of R78A-K855A (M1) and R765A-K1246A (M2) identified an "electrostatic valley," a stabilizing network where positively charged residues interact with negatively charged DNA to maintain SpCas9s structural integrity. Disrupting this valley through direct (M2) or allosteric (M1) mutations destabilized SpCas9s DNA-bound conformation, leading to distinct pathways for improving SpCas9 specificity. This study provides a new framework for understanding allostery in Cas9, integrating ML-driven structural analysis with MD simulations. By identifying key allosteric residues and introducing the electrostatic valley as a central concept, we offer a rational strategy for engineering high-fidelity Cas9 variants. Beyond Cas9, our approach can be applied to uncover allosteric hotspots in other enzyme regulation and rational protein design.

pharmacology and toxicology↗

CasGen: A Regularized Generative Model for CRISPR Cas Protein Design with Classification and Margin-Based Optimization

Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated proteins (Cas) systems have revolutionized genome editing by providing high precision and versatility. However, most genome editing applications rely on a limited number of well-characterized Cas9 and Cas12 variants, constraining the potential for broader genome engineering applications. In this study, we extensively explored Cas9 and Cas12 proteins and developed CasGen, a novel transformer-based deep generative model with margin-based latent space regularization to enhance the quality of newly generative Cas9 and Cas12 proteins. Specifically, CasGen employs a strategies that combine classification to filter out non-Cas sequences, Bayesian optimization of the latent space to guide functionally relevant designs, and thorough structural validation using AlphaFold-based analyses to ensure robust protein generation. We collected a comprehensive dataset with 3,021 Cas9, 597 Cas12, and 597 Non-Cas protein sequences from reputable biological databases such as InterPro and PDB. To validate the generated proteins, we performed sequence alignment using the BLAST tool to ensure novelty and filter out highly similar sequences to existing Cas proteins. Structural prediction using AlphaFold2 and AlphaFold3 confirmed that the generated proteins exhibit high structural similarity to known Cas9 and Cas12 variants, with TM-scores between 0.70 and 0.85 and root-mean-square deviation (RMSD) values below 2.00 [A]. Sequence identity analysis further demonstrated that the generated Cas9 orthologs exhibited 28% to 55% identity with known variants, while Cas12a variants show up to 48% identity. Our results demonstrate that the proposed Cas generative model has significant potential to expand the genome editing toolkit by designing diverse Cas proteins that retain functional integrity. The developed deep generative approach offers a promising avenue for synthetic biology and therapeutic applications, enableling the development of more precise and versatile Cas-based genome editing tools.

genomics↗