Search bioRxiv⌕ Search

Biology subjects

Consbruck, R.

Publications and source records attributed to Consbruck, R..

3 recordsLinked to original sources

In vitro validated antibody design against multiple therapeutic antigens using generative inverse folding

Deep learning approaches have demonstrated the ability to design protein sequences given backbone structures [1, 2, 3, 4, 5]. While these approaches have been applied in silico to designing antibody complementarity-determining regions (CDRs), they have yet to be validated in vitro for designing antibody binders, which is the true measure of success for antibody design. Here we describe IgDesign, a deep learning method for antibody CDR design, and demonstrate its robustness with successful binder design for 8 therapeutic antigens. The model is tasked with designing heavy chain CDR3 (HCDR3) or all three heavy chain CDRs (HCDR123) using native backbone structures of antibody-antigen complexes, along with the antigen and antibody framework (FWR) sequences as context. For each of the 8 antigens, we design 100 HCDR3s and 100 HCDR123s, scaffold them into the native antibodys variable region, and screen them for binding against the antigen using surface plasmon resonance (SPR). As a baseline, we screen 100 HCDR3s taken from the models training set and paired with the native HCDR1 and HCDR2. We observe that both HCDR3 design and HCDR123 design outperform this HCDR3-only baseline. IgDesign is the first experimentally validated antibody inverse folding model. It can design antibody binders to multiple therapeutic antigens with high success rates and, in some cases, improved affinities over clinically validated reference antibodies. Antibody inverse folding has applications to both de novo antibody design and lead optimization, making IgDesign a valuable tool for accelerating drug development and enabling therapeutic design. The data generated in this study serve as a useful benchmark of diverse antibody-antigen interactions. We use this data to benchmark self-consistency RMSD (scRMSD), using ABodyBuilder2 [6], ABodyBuilder3 [7], and ESMFold [8], as a metric for assessing binding. We open source the code for IgDesign and the SPR datasets. Code and datasets can be found at https://github.com/AbSciBio/igdesign.

synthetic biology↗

Deep learning-based codon optimization with large-scale synonymous variant datasets enables generalized tunable protein expression

Increasing recombinant protein expression is of broad interest in industrial biotechnology, synthetic biology, and basic research. Codon optimization is an important step in heterologous gene expression that can have dramatic effects on protein expression level. Several codon optimization strategies have been developed to enhance expression, but these are largely based on bulk usage of highly frequent codons in the host genome, and can produce unreliable results. Here, we develop deep contextual language models that learn the codon usage rules from natural protein coding sequences across members of the Enterobacterales order. We then fine-tune these models with over 150,000 functional expression measurements of synonymous coding sequences from three proteins to predict expression in E. coli. We find that our models recapitulate natural context-specific patterns of codon usage and can accurately predict expression levels across synonymous sequences. Finally, we show that expression predictions can generalize across proteins unseen during training, allowing for in silico design of gene sequences for optimal expression. Our approach provides a novel and reliable method for tuning gene expression with many potential applications in biotechnology and biomanufacturing.

synthetic biology↗

Antibody optimization enabled by artificial intelligence predictions of binding affinity and naturalness

Traditional antibody optimization approaches involve screening a small subset of the available sequence space, often resulting in drug candidates with suboptimal binding affinity, developability or immunogenicity. Based on two distinct antibodies, we demonstrate that deep contextual language models trained on high-throughput affinity data can quantitatively predict binding of unseen antibody sequence variants. These variants span a KD range of three orders of magnitude over a large mutational space. Our models reveal strong epistatic effects, which highlight the need for intelligent screening approaches. In addition, we introduce the modeling of "naturalness", a metric that scores antibody variants for similarity to natural immunoglobulins. We show that naturalness is associated with measures of drug developability and immunogenicity, and that it can be optimized alongside binding affinity using a genetic algorithm. This approach promises to accelerate and improve antibody engineering, and may increase the success rate in developing novel antibody and related drug candidates.

bioinformatics↗