Search bioRxiv⌕ Search

Biology subjects

Ahern, W.

Publications and source records attributed to Ahern, W..

4 recordsLinked to original sources

Computational Design of Metallohydrolases

De novo enzyme design starts from a description of an ideal active site composed of catalytic residues surrounding the reaction transition state(s), and builds a protein structure that contains this site1-7. Generative AI methods such as RFdiffusion11,12 now enable the direct generation of proteins around active sites, but to date, such scaffolding has required specification of both the position in the sequence and the backbone coordinates of each catalytic residue, which complicates sampling. Here we introduce a generative AI method called RFdiffusion2 that overcomes these limitations and use it to design zinc metallohydrolases starting from a density functional theory description of the active site geometry. Of an initial set of 96 designs tested experimentally, the most active has a kcat/KM of 16,000 M-1 s-1, orders of magnitude higher than previously designed metallohydrolases.6,7,13,14 A second round of 96 designs yielded 3 additional highly active enzymes, with kcat/KM up to 53,000 M-1 s-1 and kcat up to 1.5 s-1. The structures of the four enzymes are very different from each other and from the structures in the PDB. Each enzyme positions the reaction substrate almost perfectly for nucleophilic attack by a water molecule activated by the bound metal, and are predicted by PLACER15 and Chai-144 to have highly preorganized active sites. The ability to generate highly active catalysts straight out of the computer, without experimental optimization, using quantum chemistry calculated active site geometries should open the door to a new generation of potent designer enzymes.16,17

biochemistry↗

Computational design of serine hydrolases

Enzymes that proceed through multistep reaction mechanisms often utilize complex, polar active sites positioned with sub-angstrom precision to mediate distinct chemical steps, which makes their de novo construction extremely challenging. We sought to overcome this challenge using the classic catalytic triad and oxyanion hole of serine hydrolases as a model system. We used RFdiffusion1 to generate proteins housing catalytic sites of increasing complexity and varying geometry, and a newly developed ensemble generation method called ChemNet to assess active site geometry and preorganization at each step of the reaction. Experimental characterization revealed novel serine hydrolases that catalyze ester hydrolysis with catalytic efficiencies (kcat/Km) up to 3.8 x 103 M-1 s-1, closely match the design models (C RMSDs < 1 [A]), and have folds distinct from natural serine hydrolases. In silico selection of designs based on active site preorganization across the reaction coordinate considerably increased success rates, enabling identification of new catalysts in screens of as few as 20 designs. Our de novo buildup approach provides insight into the geometric determinants of catalysis that complements what can be obtained from structural and mutational studies of native enzymes (in which catalytic group geometry and active site makeup cannot be so systematically varied), and provides a roadmap for the design of industrially relevant serine hydrolases and, more generally, for designing complex enzymes that catalyze multi-step transformations.

biochemistry↗

Generalized Biomolecular Modeling and Design with RoseTTAFold All-Atom

Although AlphaFold2 (AF2) and RoseTTAFold (RF) have transformed structural biology by enabling high-accuracy protein structure modeling, they are unable to model covalent modifications or interactions with small molecules and other non-protein molecules that can play key roles in biological function. Here, we describe RoseTTAFold All-Atom (RFAA), a deep network capable of modeling full biological assemblies containing proteins, nucleic acids, small molecules, metals, and covalent modifications given the sequences of the polymers and the atomic bonded geometry of the small molecules and covalent modifications. Following training on structures of full biological assemblies in the Protein Data Bank (PDB), RFAA has comparable protein structure prediction accuracy to AF2, excellent performance in CAMEO for flexible backbone small molecule docking, and reasonable prediction accuracy for protein covalent modifications and assemblies of proteins with multiple nucleic acid chains and small molecules which, to our knowledge, no existing method can model simultaneously. By fine-tuning on diffusive denoising tasks, we develop RFdiffusion All-Atom (RFdiffusionAA), which generates binding pockets by directly building protein structures around small molecules and other non-protein molecules. Starting from random distributions of amino acid residues surrounding target small molecules, we design and experimentally validate proteins that bind the cardiac disease therapeutic digoxigenin, the enzymatic cofactor heme, and optically active bilin molecules with potential for expanding the range of wavelengths captured by photosynthesis. We anticipate that RFAA and RFdiffusionAA will be widely useful for modeling and designing complex biomolecular systems.

biochemistry↗

Broadly applicable and accurate protein design by integrating structure prediction networks and diffusion generative models

There has been considerable recent progress in designing new proteins using deep learning methods1-9. Despite this progress, a general deep learning framework for protein design that enables solution of a wide range of design challenges, including de novo binder design and design of higher order symmetric architectures, has yet to be described. Diffusion models10,11 have had considerable success in image and language generative modeling but limited success when applied to protein modeling, likely due to the complexity of protein backbone geometry and sequence-structure relationships. Here we show that by fine tuning the RoseTTAFold structure prediction network on protein structure denoising tasks, we obtain a generative model of protein backbones that achieves outstanding performance on unconditional and topology-constrained protein monomer design, protein binder design, symmetric oligomer design, enzyme active site scaffolding, and symmetric motif scaffolding for therapeutic and metal-binding protein design. We demonstrate the power and generality of the method, called RoseTTAFold Diffusion (RFdiffusion), by experimentally characterizing the structures and functions of hundreds of new designs. In a manner analogous to networks which produce images from user-specified inputs, RFdiffusion enables the design of diverse, complex, functional proteins from simple molecular specifications.

biochemistry↗