Search bioRxiv⌕ Search

Biology subjects

Jaakkola, T.

Publications and source records attributed to Jaakkola, T..

3 recordsLinked to original sources

Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction

Accurately modeling biomolecular interactions is a central challenge in modern biology. While recent advances, such as AlphaFold3 and Boltz-1, have substantially improved our ability to predict biomolecular complex structures, these models still fall short in predicting binding affinity, a critical property underlying molecular function and therapeutic efficacy. Here, we present Boltz-2, a new structural biology foundation model that exhibits strong performance for both structure and affinity prediction. Boltz-2 introduces controllability features including experimental method conditioning, distance constraints, and multi-chain template integration for structure prediction, and is, to our knowledge, the first AI model to approach the performance of free-energy perturbation (FEP) methods in estimating small molecule-protein binding affinity. Crucially, it achieves strong correlation with experimental readouts on many benchmarks, while being at least 1000x more computationally efficient than FEP. By coupling Boltz-2 with a generative model for small molecules, we demonstrate an effective workflow to find diverse, synthesizable, high-affinity binders, as estimated by absolute FEP simulations on the TYK2 target. To foster broad adoption and further innovation at the intersection of machine learning and biology, we are releasing Boltz-2 weights, inference, and training code 1 under a permissive open license, providing a robust and extensible foundation for both academic and industrial research.

molecular biology↗

Atom level enzyme active site scaffolding using RFdiffusion2

De novo enzyme design starts from ideal active site descriptions consisting of constellations of catalytic residue functional groups around reaction transition state(s), and seeks to generate protein structures that can accurately hold the site in place. Highly active enzymes have been designed starting from such descriptions using the generative AI method RFdiffusion [1-3], but there are two current methodological limitations. First, the geometry of the active site can only be specified at the residue level, so for each catalytic residue functional group placed around the reaction transition state, the possible locations of the residue backbone must be enumerated by building side chain rotamers back from the functional group. Second, the location of the catalytic residues along the sequence must be specified in advance, which considerably limits the space of solutions which can be sampled. Here we describe a new deep generative method, Rosetta Fold diffusion 2 (RFdiffusion2), that solves both problems, enabling enzymes to be designed from sequence agnostic descriptions of functional group locations without inverse rotamer generation. We first evaluate RFdiffusion2 on an in silico enzyme design benchmark of 41 diverse active sites and find that it is able to successfully build proteins scaffolding all 41 sites, compared to 16/41 with prior state-of-the-art deep learning methods. Next, we design enzymes around three diverse catalytic sites and characterize the designs experimentally; in each case we identify active catalysts in testing less than 96 sequences. RFdiffusion2 demonstrates the potential of atomic resolution generative models for the design of de novo enzymes directly from their reaction mechanisms.

biochemistry↗

Boltz-1: Democratizing Biomolecular Interaction Modeling

Understanding biomolecular interactions is fundamental to advancing fields like drug discovery and protein design. In this paper, we introduce BO_SCPLOWOLTZC_SCPLOW-1, an open-source deep learning model incorporating innovations in model architecture, speed optimization, and data processing achieving AO_SCPLOWLPHAC_SCPLOWFO_SCPLOWOLDC_SCPLOW3-level accuracy in predicting the 3D structures of biomolecular complexes. BO_SCPLOWOLTZC_SCPLOW-1 demonstrates a performance on-par with state-of-the-art commercial models on a range of diverse benchmarks, setting a new benchmark for commercially accessible tools in structural biology. Further, we push the boundary of capabilities of these models with BO_SCPLOWOLTZC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWSTEERINGC_SCPLOW, a new inference time steering technique that is able to fix hallucinations and non-physical predictions from the models. By releasing the training and inference code, model weights, datasets, and benchmarks under the MIT open license, we aim to foster global collaboration, accelerate discoveries, and provide a robust platform for advancing biomolecular modeling.

biophysics↗