Search bioRxiv⌕ Search

Biology subjects

Pacesa, M.

Publications and source records attributed to Pacesa, M..

8 recordsLinked to original sources

BindCraft: one-shot design of functional protein binders

Protein-protein interactions (PPIs) are at the core of all key biological processes. However, the complexity of the structural features that determine PPIs makes their design challenging. We present BindCraft, an open-source and automated pipeline for de novo protein binder design with experimental success rates of 10-100%. BindCraft leverages the weights of AlphaFold21 to generate binders with nanomolar affinity without the need for high-throughput screening or experimental optimization, even in the absence of known binding sites. We successfully designed binders against a diverse set of challenging targets, including cell-surface receptors, common allergens, de novo designed proteins, and multi-domain nucleases, such as CRISPR-Cas9. We showcase the functional and therapeutic potential of designed binders by reducing IgE binding to birch allergen in patient-derived samples, modulating Cas9 gene editing activity, and reducing the cytotoxicity of a foodborne bacterial enterotoxin. Lastly, we utilize cell surface receptor-specific binders to redirect AAV capsids for targeted gene delivery. This work represents a significant advancement towards a "one design-one binder" approach in computational design, with immense potential in therapeutics, diagnostics, and biotechnology.

bioinformatics↗

Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences

Gene editing has the potential to solve fundamental challenges in agriculture, biotechnology, and human health. CRISPR-based gene editors derived from microbes, while powerful, often show significant functional tradeoffs when ported into non-native environments, such as human cells. Artificial intelligence (AI) enabled design provides a powerful alternative with potential to bypass evolutionary constraints and generate editors with optimal properties. Here, using large language models (LLMs) trained on biological diversity at scale, we demonstrate the first successful precision editing of the human genome with a programmable gene editor designed with AI. To achieve this goal, we curated a dataset of over one million CRISPR operons through systematic mining of 26 terabases of assembled genomes and meta-genomes. We demonstrate the capacity of our models by generating 4.8x the number of protein clusters across CRISPR-Cas families found in nature and tailoring single-guide RNA sequences for Cas9-like effector proteins. Several of the generated gene editors show comparable or improved activity and specificity relative to SpCas9, the prototypical gene editing effector, while being 400 mutations away in sequence. Finally, we demonstrate an AI-generated gene editor, denoted as OpenCRISPR-1, exhibits compatibility with base editing. We release OpenCRISPR-1 publicly to facilitate broad, ethical usage across research and commercial applications.

synthetic biology↗

Targeting protein-ligand neosurfaces using a generalizable deep learning approach

Molecular recognition events between proteins drive biological processes in living systems. However, higher levels of mechanistic regulation have emerged, where protein-protein interactions are conditioned to small molecules. Here, we present a computational strategy for the design of proteins that target neosurfaces, i.e. surfaces arising from protein-ligand complexes. To do so, we leveraged a deep learning approach based on learned molecular surface representations and experimentally validated binders against three drug-bound protein complexes. Remarkably, surface fingerprints trained only on proteins can be applied to neosurfaces emerging from small molecules, serving as a powerful demonstration of generalizability that is uncommon in deep learning approaches. The designed chemically-induced protein interactions hold the potential to expand the sensing repertoire and the assembly of new synthetic pathways in engineered cells.

biochemistry↗

An atlas of protein homo-oligomerization across domains of life

Protein structures are essential to understand cellular processes in molecular detail. While advances in AI revealed the tertiary structure of proteins at scale, their quaternary structure remains mostly unknown. Here, we describe a scalable strategy based on AlphaFold2 to predict homo-oligomeric assemblies across four proteomes spanning the tree of life. We find that 50% of archaeal, 45% of bacterial, and 20% of eukaryotic proteomes form homomers. Our predictions accurately capture protein homo-oligomerization, recapitulate megadalton complexes, and unveil hundreds of novel homo-oligomer types. Analyzing these datasets reveals coiled-coil regions as major enablers of quaternary structure evolution in Eukaryotes. Integrating these structures with omics data shows that a majority of known protein complexes are symmetric. Finally, these datasets provide a structural context for interpreting disease mutations, which we find enriched at interfaces. Our strategy is applicable to any organism and provides a comprehensive view of homo-oligomerization in proteomes, protein networks, and disease. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=193 SRC="FIGDIR/small/544317v1_ufig1.gif" ALT="Figure 1"> View larger version (79K): org.highwire.dtl.DTLVardef@1507a12org.highwire.dtl.DTLVardef@7e522aorg.highwire.dtl.DTLVardef@1445410org.highwire.dtl.DTLVardef@eb09f7_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Computational design of soluble analogues of integral membrane protein structures

De novo design of complex protein folds using solely computational means remains a significant challenge. Here, we use a robust deep learning pipeline to design complex folds and soluble analogues of integral membrane proteins. Unique membrane topologies, such as those from GPCRs, are not found in the soluble proteome and we demonstrate that their structural features can be recapitulated in solution. Biophysical analyses reveal high thermal stability of the designs and experimental structures show remarkable design accuracy. The soluble analogues were functionalized with native structural motifs, standing as a proof-of-concept for bringing membrane protein functions to the soluble proteome, potentially enabling new approaches in drug discovery. In summary, we designed complex protein topologies and enriched them with functionalities from membrane proteins, with high experimental success rates, leading to a de facto expansion of the functional soluble fold space.

bioinformatics↗

De novo design of site-specific protein interactions with learned surface fingerprints

Physical interactions between proteins are essential for most biological processes governing life. However, the molecular determinants of such interactions have been challenging to understand, even as genomic, proteomic, and structural data grows. This knowledge gap has been a major obstacle for the comprehensive understanding of cellular protein-protein interaction (PPI) networks and for the de novo design of protein binders that are crucial for synthetic biology and translational applications. We exploit a geometric deep learning framework operating on protein surfaces that generates fingerprints to describe geometric and chemical features critical to drive PPIs. We hypothesized these fingerprints capture the key aspects of molecular recognition that represent a new paradigm in the computational design of novel protein interactions. As a proof-of-principle, we computationally designed several de novo protein binders to engage four protein targets: SARS-CoV-2 spike, PD-1, PD-L1, and CTLA-4. Several designs were experimentally optimized while others were purely generated in silico, reaching nanomolar affinity with structural and mutational characterization showing highly accurate predictions. Overall, our surface-centric approach captures the physical and chemical determinants of molecular recognition, enabling a novel approach for the de novo design of protein interactions and, more broadly, of artificial proteins with function.

bioengineering↗

Structural basis for Cas9 off-target activity

The target DNA specificity of the CRISPR-associated genome editor nuclease Cas9 is determined by complementarity to a 20-nucleotide segment in its guide RNA. However, Cas9 can bind and cleave partially complementary off-target sequences, which raises safety concerns for its use in clinical applications. Here we report crystallographic structures of Cas9 bound to bona fide off-target substrates, revealing that off-target binding is enabled by a range of non-canonical base pairing interactions within the guide-off-target heteroduplex. Off-target sites containing single-nucleotide deletions relative to the guide RNA are accommodated by base skipping or multiple non-canonical base pairs rather than RNA bulge formation. Additionally, PAM-distal mismatches result in duplex unpairing and induce a conformational change of the Cas9 REC lobe that perturbs its conformational activation. Together, these insights provide a structural rationale for the off-target activity of Cas9 and contribute to the improved rational design of guide RNAs and off-target prediction algorithms.

biochemistry↗

Mechanism of R-loop formation and conformational activation of Cas9

Cas9 is a CRISPR-associated endonuclease capable of RNA-guided, site-specific DNA cleavage1-3. The programmable activity of Cas9 has been widely utilized for genome editing applications4-6, yet its precise mechanisms of target DNA binding and off-target discrimination remain incompletely understood. Here we report a series of cryo-EM structures of Streptococcus pyogenes Cas9 capturing the directional process of target DNA hybridization. In the early phase of R-loop formation, the Cas9 REC2 and REC3 domains form a positively charged cleft that accommodates the distal end of the target DNA duplex. Guide-target hybridization past the seed region induces rearrangements of the REC2/3 domains and relocation of the HNH nuclease domain to assume a catalytically incompetent checkpoint conformation. Completion of the guide-target heteroduplex triggers conformational activation of the HNH nuclease domain, enabled by distortion of the guide-target heteroduplex, and complementary REC2/3 domain rearrangements. Together, these results establish a structural framework for target DNA-dependent activation of Cas9 that sheds light on its conformational checkpoint mechanism and may facilitate the development of novel Cas9 variants and guide RNA designs with enhanced specificity and activity.

biochemistry↗