Search bioRxiv⌕ Search

Biology subjects

Diethe, T.

Publications and source records attributed to Diethe, T..

2 recordsLinked to original sources

Generative Language Modeling for Antibody CDR Grafting and Alignment-driven De Novo Design

Antibodies recognise their targets through hypervariable complementarity-determining regions (CDRs), which are interleaved with conserved frameworks in sequence space, making de novo CDR design an infilling problem. Autoregressive models generate residues left-to-right, which precludes full framework context during CDR generation and conflates framework and CDR likelihoods, leaving no natural prompt-response interface for feedback to steer generation. We present GenCDR, a family of LLaMa-based autoregressive language models that read all frameworks as a conditioning prompt and generate all CDRs jointly as a variable-length response, making CDR likelihoods a clean, separable target for reward attribution. The family comprises IgGenCDR, p-IgGenCDR, and NanoGenCDR, trained on unpaired, paired, and nanobody chains, respectively. GenCDR achieves the highest CDR recovery among autoregressive models and produces natural, diverse, human-like CDRs whose likelihoods correlate with fitness and developability assays. The prompt-response boundary also enables principled alignment: reward signals for binding affinity, expression, or developability can be composed to steer CDR generation. Over four rounds of alignment against antibody-antigen co-folding and developability objectives, we find that NanoGenCDR, which uses no explicit antigen encoding, can reach in silico structural interface metrics competitive with those of a structure-conditioned diffusion pipeline at roughly half the sampling budget, with more natural, developable designs. The same interface can be extended to integrate experimental feedback, opening a path to closed-loop antibody de novo design.

synthetic biology↗

Representation Learning of Human Disease Mechanisms for a Foundation Model in Rare and Common Diseases

A fundamental challenge in translational medicine is the computational modeling of complex human diseases to accelerate therapeutic development. Representation learning provides a powerful framework to address this, yet creating models that capture deep biological mechanisms remains a critical need. To this end, we propose a novel strategy that partitions the disease landscape into rare and non-rare categories, enabling systematic knowledge repurposing both within and between these groups. Here, we introduce Dis2Vec (Disease to Vector), a representation learning framework designed to operationalize this concept. Dis2Vec generates biologically grounded disease embeddings by learning from human genetic and phenotypic data, forming the foundation for Disease-Disease Association Learning (DDAL) and unsupervised disease clustering. We evaluate Dis2Vec representations in two downstream applications. First, we assess DDAL performance on a transfer learning benchmark designed to predict therapeutic transferability, using real-world drug repurposing investment decisions made in clinical trials. Second, unsupervised clustering analyses reveal shared biological mechanisms across diseases. By modeling the disease landscape in this way, Dis2Vec enhances translational research efficiency across both rare and non-rare diseases, advancing the development of foundational models for therapeutic science. Furthermore, Dis2Vec establishes a biologically grounded disease-representation and benchmarking layer that paves the way for trustworthy agentic biomedical AI systems in rare-disease indication expansion.

systems biology↗