Search bioRxiv⌕ Search

Biology subjects

Liang, W.-C.

Publications and source records attributed to Liang, W.-C..

4 recordsLinked to original sources

Property Enhancer - a data efficient multi-objective approach for functional antibody optimization

In-silico antibody lead optimization remains challenging due to scarce high-quality data, costly experimental validation, and the need to jointly optimize multiple developability properties. Discovery workflows often rely on high-throughput phage, ribosome or yeast display experiments, which yield large but noisy datasets; as leads emerge, strategies shift to low-throughput assays which are precise, yet unscalable. Deep-learning and language-model approaches are hindered by such limited, unreliable measurements. We introduce Property Enhancer (PropEn), a data-efficient framework for low-data, heterogeneous regimes that can simultaneously optimize multiple antibody properties. PropEn proposes a matching-based augmentation that expands the training data with sequence pairs differing by only a few mutations; within each pair the second sequence improves the target value, providing an implicit optimization signal. Extensive in silico and in vitro tests show 10-39x affinity gains across four targets and nine leads, and enable joint multi-property optimization, positioning PropEn as a scalable, general solution.

molecular biology↗

Rational design of oxidation-resistant antibodies through local electrostatic modulation

Oxidation is a significant degradation pathway in proteins, particularly therapeutic antibodies, that can impair function, efficacy, and stability. Understanding the sequence and structural factors that drive oxidation susceptibility is critical for incorporating chemical stability into early-stage antibody design. Here, we present a machine learning classifier trained on expert-guided structural features to assess tryptophan (Trp) oxidation risks in the complementarity-determining regions (CDR) of 187 antibodies produced internally at Genentech. The model reveals a strong correlation between local negative electrostatic potential and Trp oxidation susceptibility. A simplified two-parameter model derived from these insights achieves 79% accuracy in classifying oxidation risk, compared to 84% accuracy with the full-feature classifier. Beyond our internal dataset, the two-parameter model successfully predicted oxidation risk for all CDR Trp sites in a blind subset of eight clinical-stage antibodies. In addition, we show that modulating the electrostatic potential around the Trp side-chains through distant charge-altering mutations can significantly reduce oxidation rates. In four out of five re-engineered antibodies, oxidation rates decreased by approximately 50%, while half of these maintained binding affinity. Finally, we demonstrate that this approach can guide multi-property optimization by balancing oxidation resistance and affinity in an anti-CD33 antibody. These results establish a strong link between local electrostatic environments and Trp oxidation susceptibility, and provide a practical framework for designing oxidation resistant biotherapeutics.

biochemistry↗

Lab-in-the-loop therapeutic antibody design with deep learning

Therapeutic antibody design is a complex multi-property optimization problem with substantial promise for improvement with the application of machine-learning methods. Towards realizing that promise, we introduce "Lab-in-the-loop," a new approach that orchestrates state-of-the-art repertoire mining methods, generative machine learning models, multi-task property predictors, active learning ranking and selection, and in vitro experimentation in a semi-autonomous, iterative optimization loop. By automating the design of antibody variants, property prediction, ranking and selection of designs to assay in the lab, and ingestion of in vitro data, we enable an end-to-end approach to developing computationally-informed therapeutic antibody design pipelines. We apply lab-in-the-loop to eleven seed antibodies obtained via animal immunization with four clinically relevant antigen targets: EGFR, IL-6, HER2, and OSM. Over 1,800 unique antibody variants are tested throughout four rounds of iterative optimization identifying 3-100x better binding variants for all targets and 10/11 seeds, with the best binders exceeding 100 pM affinity, demonstrating a process by which end-to-end machine learning can be developed for therapeutic antibody development.

bioengineering↗

DyAb: sequence-based antibody design and property prediction in a low-data regime

Protein therapeutic design and property prediction are frequently hampered by data scarcity. Here we propose a new model, DyAb, that addresses these issues by leveraging a pair-wise representation to predict differences in protein properties, rather than absolute values. DyAb is built on top of a pre-trained protein language model and achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across molecules targeting three different antigens (EGFR, IL-6, and an internal target), given as few as 100 training data. We employ DyAb in two design contexts: as a ranking model to score combinations of known mutations, and combined with a genetic algorithm to generate new sequences. Our method consistently generates novel antibody candidates with high binding rates, including designs that improve on the binding affinity of the lead molecule by more than ten-fold. DyAb represents a powerful tool for engineering therapeutic protein properties in low data regimes common in early-stage drug development.

bioengineering↗