Search bioRxiv⌕ Search

Biology subjects

Chomicz, D.

Publications and source records attributed to Chomicz, D..

6 recordsLinked to original sources

NAStructuralDB : Structural database to facilitate computational studies of molecular modeling and recognition of proteins with special focus on antibody-antigen interactions.

Studying the interactions between antibodies and antigens is fundamental to the development of novel therapeutic biologics. Predictions of such interactions start with data collection. Though there exist reliable resources to identify antibody structures in the Protein Data Bank (PDB), such data still requires substantial processing to be usable in predictive tasks. Redundancy in sequences needs to be removed to avoid data leakages between train, test and validation sets. Descriptors such as surface accessibility, secondary structure and antibody region information need to be additionally annotated. Information on inter- and intra-molecular contacts, which is crucial to studying paratope/epitope information, needs to be collected. The specialized immunoglobulin format of Nanobodies(R) requires a separate dataset mirroring that of antibodies, given that their structure contains only a single VHH chain. Because antibody-antigen structures account for a small amount of all protein-protein contacts, having a molecular contact reference from other proteins is also desired. To address these issues, we introduce NAStructuralDB (https://naturalantibody.com/na-structural/), a dataset of processed structures of antibodies, Nanobodies(R), proteins and their complexes with molecular contact information and associated annotations. We use the opportunity of having collected the contact data to provide a reference of binding propensities of different residues across distinct contact types. We anticipate that this dataset will accelerate a broad range of predictive tasks by standardizing common, time-consuming data preparation steps in antibody and protein design.

bioinformatics↗

Benchmarking antigen-aware inverse folding methods for antibody design.

Computational antibody design has seen many recent advances pioneered via the use of language models and advanced structure prediction tools. Developing a de novo antibody against a specific antigen requires structural awareness that most language models lack. A prominent class of machine learning methods combining the best of language model and structural worlds is inverse folding. This approach aims to predict a sequence that would fit a given structure. Such methods are now increasingly used to predict alternate sequences given a structure of a binder. It is known that, just like language models, such methods have certain predictive power in identifying binders. Here we performed a set of tests to reveal where, if at all, such methods provide value in the realistic setting of antibody discovery.

bioinformatics↗

AbDesign: Database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models.

Antibodies are naturally evolved molecular recognition scaffolds that can bind a variety of surfaces. Their designability is crucial to the development of biologics, with computational methods holding promise in accelerating the delivery of medicines to the clinic. Modeling antibody-antigen recognition is prohibitively difficult, with data paucity being one of the biggest hurdles. Current affinity datasets comprise a small number of experimental measurements, which are often not standardized between molecules. Here, we address these issues by creating a dataset of seven antigens with two antibodies each, for which we introduce a heterogeneous set of mutations to the CDR-H3 measured by ELISA. Each of the parental complexes has a known crystal structure. We perform benchmarking of state-of-the-art affinity prediction algorithms to gauge their effectiveness. Current computational methods exhibit significant limitations in accurately predicting the effects of single-point mutations. In contrast, the older empirical, physics-based method FoldX, performs well in identifying mutants that retain binding. These findings highlight the need for more resources like the one presented here -- large, molecularly diverse, and experimentally consistent datasets.

bioinformatics↗

nanoFOLD : sequence design of nanobodies via inverse folding

Antibodies devoid of light chains are a promising class of biotherapeutics. Computational methods that address these molecules are crucially needed to accelerate the traditional, long and expensive experimental process of their discovery. Inverse folding, wherein one is tasked to predict a sequence given molecular coordinates, is an established method in scaffold-based protein design. Here we develop an inverse folding method speci[fi]c to nanobodies. We demonstrate its application in nanobody-engineering scenarios of enriching binders from next-generation sequencing experiments and novel binder design.

bioinformatics↗

Conserved heavy/light contacts and germline preferences revealed by a large-scale analysis of natively paired human antibody sequences and structural data.

Antibody next-generation sequencing (NGS) datasets have become crucial to develop computational models addressing this successful class of therapeutics. Although antibodies are composed of both heavy and light chains, most NGS sequencing depositions provide them in unpaired form, reducing their utility. Here we introduce PairedAbNGS, a novel database with paired heavy/light antibody chains. To the best of our knowledge, this is the largest resource for paired natural antibody sequences with 58 bioprojects and over 14 million assembled productive sequences. We make the database accessible at http://naturalantibody.com/paired-ngs as a valuable tool for biological and machine-learning applications. Using this dataset, we investigated heavy and light chain variable (V) gene pairing preferences and found significant biases beyond gene usage frequencies, possibly due to receptor editing favoring less autoreactive combinations. Analyzing the available antibody structures from the Protein Data Bank, we studied conserved contact residues between heavy and light chains, particularly interactions between the CDR3 region of one chain and the FWR2 region of the opposite chain. Examination of amino acid pairs at key contact sites revealed significant deviations of amino acids distributions compared to random pairings, in the heavy chains CDR3 region contacting the opposite chain, indicating specific interactions might be crucial for proper chain pairing. This observation is further reinforced by preferential IGHV-IGLJ and IGLV-IGHJ pairing preferences. We hope that both our resources and the findings would contribute to improving the engineering of biological drugs.

immunology↗

RIOT - Rapid Immunoglobulin Overview Tool - annotation of nucleotide and amino acid immunoglobulin sequences using an open germline database.

Antibodies are a cornerstone of the immune system, playing a pivotal role in identifying and neutralizing infections caused by bacteria, viruses, and other pathogens. Understanding their structure, and function, can provide insights into both the bodys natural defenses and the principles behind many therapeutic interventions, including vaccines and antibody-based drugs. The analysis and annotation of antibody sequences, including the identification of variable, diversity, joining, and constant genes, as well as the delineation of framework regions and complementarity-determining regions, is essential for understanding their structure and function. Currently analyzing large volumes of antibody sequences is routine in antibody discovery, requiring fast and accurate tools. While there are existing tools designed for the annotation and numbering of antibody sequences, they often have limitations such as being restricted to either nucleotide or amino acid sequences, reliance on non-uniform germline databases, or slow execution times. Here we present Rapid Immunoglobulin Overview Tool (RIOT), a novel open-source solution for antibody numbering that addresses these shortcomings. RIOT handles nucleotide and amino acid sequence processing, comes with a free germline database, and is computationally efficient. We hope the tool will facilitate rapid annotation of antibody sequencing outputs for the benefit of understanding antibody biology and discovering novel therapeutics. AvailabilityRIOT is available at https://github.com/NaturalAntibody/riot_na.

bioinformatics↗