Search bioRxiv⌕ Search

Biology subjects

Chungyoun, M.

Publications and source records attributed to Chungyoun, M..

5 recordsLinked to original sources

ProtAff: Protein Binding Affinity Prediction via LoRA-Finetuned ESM-2

Predicting the binding affinity of protein-protein interactions remains a central challenge in computational biology. Structure prediction models such as AlphaFold3 (AF3) and Boltz-2 can produce high-quality docking poses, and their confidence scores indicate structure quality, but these same scores fail to rank binding affinity among confirmed binders. Here we present ProtAff, a sequence-only affinity prediction model built on ESM-2 (650M parameters) with low-rank adaptation (LoRA) fine-tuning and a cross-attention module. ProtAff is trained using a margin ranking loss on 362,567 affinity measurements spanning 20 heterogeneous data sources, and we removed all training samples whose target sequence exceeds 50% similarity to the test target EGFR. On the AdaptyvBio EGFR benchmark (N =55), ProtAff achieves a Spearman correlation coefficient{rho} = 0.413, outperforming the best AF3 metric ({rho} = 0.054), the best Boltz-2 metric ({rho} = -0.046), and ML-based predictors MINT ({rho} = 0.242) and CrossAffinity ({rho} = 0.216). Applied to the AdaptyvBio Nipah virus binder design competition, a pipeline incorporating ProtAff for affinity ranking produced a design with KD = 0.132 nM (2 of 5 designs confirmed binding), a 2.8-fold improvement over the competition winner. On a cross-target discrimination benchmark of 91 VHH-antigen crystal structures, ProtAff underperforms structural methods for distinguishing cognate from non-cognate pairings, indicating that sequence-based affinity models are effective for within-target ranking but not for cross-target specificity.

bioinformatics↗

GermRL: Alleviating The Germline Bias In Autoregressive Antibody Language Models Through Reinforcement Learning

Antibodies are powerful therapeutics whose antigen specificity arises from sequence diversity shaped during development. Recently, language models trained on large antibody repertoire datasets have enabled the generation and screening of novel candidates, but these models retain a strong germline bias. As AI adoption increases in therapeutic workflows, it is crucial to develop models that harness the diversity of antibodies necessary for the discovery of mutations that encode desirable properties. Previous work explored the germline bias in masked antibody language models, yet the bias in generative autoregressive language models has not yet been addressed. Here, we present GermRL, a lightweight and modular reinforcement learning (RL) framework capable of alleviating the germline bias in pre-trained antibody autoregressive language models through group relative policy optimization (GRPO). GermRL achieves consistent one-shot generation of antibodies that satisfy specified mutation thresholds from germline while maintaining structural plausibility. Under the lowest and highest mutation thresholds tested (5 and 35 mutations from germline), GermRL scores 0.992 and 0.950 pass@1, respectively, compared to 0.398 and 0.034 for the pre-trained language model. Within GermRL, we introduce a key pair of modifications to GRPO that increase training efficiency by discouraging reward hacking under our antibody application. Furthermore, comparison of RL generated and natural antibody sequences reveals how RL based optimization can explore alternative evolutionary mutational patterns and residue compositional strategies while preserving key global properties of natural antibodies, including identifiable germline assignments, embedding-level similarity and comparable developability profiles. Thus, RL-trained generative models optimized to promote antibody mutations through diversity from germline provide a promising framework for navigating the antibody sequence landscape, enabling exploration of novel yet biologically plausible candidates for therapeutic design.

bioinformatics↗

Fitness Landscape for Antibodies 2: Benchmarking Reveals That Protein AI Models Cannot Yet Consistently Predict Developability Properties

1A prominent application of machine learning in therapeutic antibody design is the development of models that can generate or screen antibody candidates with a high probability of success in manufacturing and clinical trials. These models must accurately represent sequence-structure-function relationships, also known as the fitness landscape. Previous protein function benchmarks examine fitness landscapes across diverse protein families, but they exclude antibody data. Here, we introduce the second iteration of the Fitness Landscape for Antibodies (FLAb2), the largest public therapeutic antibody design benchmark to date. The datasets collected in FLAb2 contain developability assay data for over 4M antibodies across 32 studies, encompassing seven properties of therapeutic antibodies: thermostability, expression, aggregation, binding affinity, pharmacokinetics, polyreactivity, and immunogenicity. Using the curated data, we evaluate the performance of 30 artificial intelligence (AI) and biophysical models in learning these properties. Protein AI models on average do not produce statistically significant correlations for most (80%) of developability datasets. No models correlate with all properties or across multiple datasets of similar properties. Zero-shot predictions from pretrained models are incapable of accurately predicting all developability properties, although several models (IgLM, ProGen2, Chai-1, ESM2, ISM, IgFold) produce statistically significant correlations for multiple datasets for thermostability, expression, binding, or immunogenicity. Fine-tuning with at least 102 points improves performance on thermostability, aggregation, and binding, but polyreactivity and pharmacokinetics lack enough data for significance. Yet it is humbling to observe that given enough developability data (103 points), a fine-tuned one-hot encoding model can match the performance of fine-tuned billion-parameter pretrained models. Training data composition influences performance more than model architecture, and intrinsic biophysical properties (thermostability) are more readily learned than extrinsic properties (immunogenicity, pharmacokinetics). Controlling for germline distance with partial correlation reveals that protein language models draw substantially on evolutionary signal; on average, germline edit distance accounts for 40% of their apparent predictive power. FLAb2 data are accessible at https://github.com/Graylab/FLAb, together with scripts that allow researchers to benchmark, compare, and iteratively improve new AI-based developability prediction models.

bioinformatics↗

Anti-citrullinated protein antibodies arise during affinity maturation of germline antibodies to carbamylated proteins in rheumatoid arthritis

Why autoantibodies in rheumatoid arthritis (RA) primarily target physiologically modified proteins, called citrullinated proteins, is unknown. Recognizing the inciting event in the production of anti-citrullinated protein antibodies (ACPAs) may shed light on the origin of RA. Here, we demonstrate that ACPAs originate from germline-encoded antibodies targeting a distinct but structurally similar modification, called carbamylation, which is pathogenic and environmentally driven. The transition from anti-carbamylated protein (anti-CarP) antibodies to ACPAs results from somatic hypermutations, indicating that the change in reactivity is acquired via antigen-driven affinity maturation. During this process, a single germline anti-CarP antibody transitions from anti-CarP to double positive (anti-CarP/ACPA) to ACPA according to the pattern and number of somatic hypermutations, explaining their coexistence and diverse specificity in RA. Artificial intelligence-based structural modeling revealed that an ACPA and its germline precursor exhibit distinct structural and biophysical properties, and pointed to heavy-chain tryptophan 48 (H-W48) as a critical residue in the differential recognition of citrullinated vs. carbamylated proteins. Indeed, a single methionine substitution in H-W48 changes the antibody specificity from ACPA to anti-CarP. These data indicate that the existence of germline-encoded anti-CarP antibodies is most likely the first event in the production of ACPAs during the early stages of RA development.

immunology↗

FLAb: Benchmarking deep learning methods for antibody fitness prediction

The successful application of machine learning in therapeutic antibody design relies heavily on the ability of models to accurately represent the sequence-structure-function landscape, also known as the fitness landscape. Previous protein bench-marks (including The Critical Assessment of Function Annotation [33], Tasks Assessing Protein Embeddings [23], and FLIP [6]) examine fitness and mutational landscapes across many protein families, but they either exclude antibody data or use very little of it. In light of this, we present the Fitness Landscape for Antibodies (FLAb), the largest therapeutic antibody design benchmark to date. FLAb currently encompasses six properties of therapeutic antibodies: (1) expression, (2) thermosta-bility, (3) immunogenicity, (4) aggregation, (5) polyreactivity, and (6) binding affinity. We use FLAb to assess the performance of various widely adopted, pretrained, deep learning models for proteins (IgLM [28], AntiBERTy [26], ProtGPT2 [11], ProGen2 [21], ProteinMPNN [7], and ESM-IF [13]); and compare them to physics-based Rosetta [1]. Overall, no models are able to correlate with all properties or across multiple datasets of similar properties, indicating that more work is needed in prediction of antibody fitness. Additionally, we elucidate how wild type origin, deep learning architecture, training data composition, parameter size, and evolutionary signal affect performance, and we identify which fitness landscapes are more readily captured by each protein model. To promote an expansion on therapeutic antibody design benchmarking, all FLAb data are freely accessible and open for additional contribution at https://github.com/Graylab/FLAb.

bioinformatics↗