Search bioRxiv⌕ Search

Biology subjects

Duroux, P.

Publications and source records attributed to Duroux, P..

6 recordsLinked to original sources

Novel Genes and Polymorphisms in Human Immunoglobulin Light Chains Across Diverse Populations Through Comprehensive IMGT(R) Analysis

The human immunoglobulin light chain loci, kappa (IGK) and lambda (IGL), are structurally complex genomic regions with germline gene content that is not yet fully characterized. These loci are marked by extensive gene duplication, allelic diversity, and segmental duplications, features that contribute critically to the adaptive immune response. In this study, we present a comprehensive IMGT annotation of IGK and IGL using two high-quality human reference assemblies (GRCh38 and T2T-CHM13) along with 142 and 125 additional chromosomal-level haploid assemblies, respectively for each locus, from individuals representing all major human superpopulations. Detailed gene and allele annotation of the reference assemblies led to the identification of 5 novel IGKV genes and 8 new IGKV alleles, 16 new IGLV genes, and 22 novel IGLV alleles. These were confirmed through assembly read validation, presence in whole genome sequencing datasets, and recurrence in multiple assemblies. Gene-level identification across the broader dataset enabled assessment of structural variation (SV) at both loci. IGL displayed high conservation, with recurrent absence observed for only one gene. In contrast, IGK exhibited greater variability, including complete loss of the distal region in certain assemblies. This structural diversity was analyzed across superpopulations, allowing us to map potential patterns of gene presence and absence across different ancestral groups. All newly identified genes were consistently observed across individuals and genomic backgrounds. This work enhances the structural resolution of the IGK and IGL loci and expands the IMGT reference directory with newly described germline genes and alleles. The results provide a more complete view of light chain genomic diversity and serve as a valuable resource for studies of antibody gene repertoires, immunogenetic variation, monoclonal antibody development and population-level diversity.

genomics↗

Therapeutic Monoclonal Antibodies Repurposing in Oncology via IMGT/mAb-KG Embeddings

BackgroundCancer remains one of the leading causes of mortality world-wide, accounting for approximately 9.7 million deaths in 2022. Faced with this significant public health challenge, therapeutic monoclonal antibodies (mAbs) have emerged as promising alternatives that may minimize the side effects associated with conventional treatments such as radiotherapy and chemotherapy. To support mAb research and development, IMGT(R), the international ImMuno-GeneTics information system, has established two standardized data sources namely IMGT/mAb-DB, a comprehensive database for mAbs, and, more recently, IMGT/mAb-KG, a dedicated knowledge graph for mAbs. Despite these advances, the development of therapeutic mAbs remains both time-consuming and financially burdensome--costs can reach up to $2.8 billion. To address this challenge and accelerate cancer treatment, mAb repurposing represents a promising alternative. ResultsIn this study, we leveraged a subset of IMGT/mAb-KG, dedicated to the oncology domain, to develop a scientific hypothesis generation application for mAb repurposing. This application, based on knowledge graph embedding techniques, is designed to suggest potential mAb candidates for novel oncology applications. A user-friendly web interface provides access to the tool, incorporating visual support to facilitate the interpretation of generated hypotheses. This application is a decision support tool aiming to accelerate the discovery of new therapeutic applications for existing mAbs. ConclusionOur application demonstrates the potential of knowledge graph embedding techniques in the oncology domain by enabling the repurposing of existing mAbs for new therapeutic uses. Using this tool, we have identified two novel mAbs, loncastuximab tesirine and glofitamab, both currently undergoing clinical trials for the treatment of chronic lymphocytic leukemia. This decision-support tool thus facilitates the discovery of new therapeutic opportunities by effectively repositioning existing mAbs for oncological indications, potentially accelerating the development of cancer therapies and addressing critical public health needs.

bioinformatics↗

Identification of Engineered IMGT Fc Variants in IMGT/mAb-DB Therapeutic Antibodies and Fusion proteins

Monoclonal antibodies (mAbs) and fusion proteins for immune applications (FPIA) play a crucial role in treating autoimmune diseases and cancers by targeting cell-surface proteins and triggering multiple immune mechanisms. These functions are mediated by the fragment crystallizable (Fc) region of mAbs and fusion proteins, whose interaction with Fc gamma receptors (Fc{gamma}Rs) can be modulated through Fc amino acid (AA) engineering. To address this, we developed the IMGT/FcVariantsExplorer tool (https://www.imgt.org/fcvariantsexplorer/) to identify AA changes within the Fc region in mAb and fusion proteins sequences from IMGT/2Dstructure-DB, the AA sequence database of IMGT(R), the international ImMunoGeneTics information system(R). We used the IMGT(R) nomenclature of engineered Fc variants involved in antibody effector properties and formats, applying a standardized classification in five categories: Effector, Half-life, Physicochemical properties, Structure, and Hybrid. We analyzed sequences of 1,107 mAbs and fusion proteins, identifying 483 entries with Fc AA changes, resulting in 211 unique Fc variants in the dataset. We also used web scraping to retrieve associated biological data from literature. All data have been integrated into IMGT/mAb-DB, with links to sequences in IMGT/2Dstructure-DB, enabling users to query Fc variants by their Category or Effect. This curated dataset reveals key trends in antibody engineering.

bioinformatics↗

IMGT(R) at scale: FAIR, Dynamic and Automated Tools for Immune Locus Analysis

IMGT(R), the international ImMunoGeneTics information system(R), has advanced its comprehensive platform for the analysis of immunoglobulin (IG) and T cell receptor (TR) genes through the development of new automated and scalable tools. This article presents major updates aligned with IMGTs three axes of research. Axis I introduces dynamic resources such as IMGT/GeneTables, IMGT/AssemblyComparison, and IMGT/StatAssembly, enabling real-time access to annotated genomic data and quality assessment of assemblies. Axis II enhances repertoire analysis with a redesigned IMGT/GeneFrequency tool and new customization features in IMGT/V-QUEST, supporting flexible exploration of IG and TR gene expression. Axis III improves the accurate prediction of peptide-MHC thanks to IMGT/RobustpMHC. Additionally, the IMGT Knowledge Graph (IMGT-KG) and its therapeutic extension, IMGT/mAb-KG, provide semantically structured access to more than 100 million immunogenetic triplets, integrating IMGT databases and linking IMGT content to external biomedical resources. These developments promote standardization, interoperability, and integrative analysis across immunogenetics and clinical applications, reinforcing IMGTs role as a core reference in the era of FAIR data and personalized medicine.

bioinformatics↗

IMGT(R) Analysis of the Human IGH Locus: Unveiling Novel Polymorphisms and Copy Number Variations in Genome Assemblies from Diverse Ancestral Backgrounds

Unraveling the genetic complexity of the human immunoglobulin heavy chain (IGH) locus provides valuable insights into the mechanisms underlying the efficacy and specificity of the adaptive immune response. Despite its crucial role, the IGH locus remains insufficiently characterized, with its allelic diversity and polymorphisms inadequately investigated. In this study, we present an analysis of the human IGH locus, incorporating 15 human genome assemblies from diverse ancestries, including African, European, Asian, Saudi, and mixed backgrounds. Through our examination of both maternal and paternal assemblies, we uncover novel IGH alleles, copy number variations (CNV), and polymorphisms, particularly within the variable (IGHV) region. Our findings reveal extensive and previously uncharacterized genetic variability in the constant (IGHC) region and distinct IMGT CNV forms across individuals. This research contributes to a significant enrichment of the IMGT(R) IGH reference directory, databases, tools and web resources and lays the groundwork for a comprehensive IMGT(R) haplotype database which can be progressively enriched to support future studies in population-specific immune profiles and adaptive immune related disease susceptibility, as comprehensive datasets become available. Such a resource promises to propel personalized immunogenomics forward, with exciting applications in cancer immunotherapy, COVID-19, and other immune-related diseases. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=82 SRC="FIGDIR/small/665479v1_ufig1.gif" ALT="Figure 1"> View larger version (42K): org.highwire.dtl.DTLVardef@17afc2corg.highwire.dtl.DTLVardef@141d785org.highwire.dtl.DTLVardef@1ac7500org.highwire.dtl.DTLVardef@13558fc_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

IMGT/RobustpMHC: Robust Training for class-I MHCPeptide Binding Prediction

The accurate prediction of peptide-MHC class I binding probabilities is a critical endeavor in immunoinformatics, with broad implications for vaccine development and immunotherapies. While recent deep neural network based approaches have showcased promise in peptide-MHC prediction, they have two shortcomings: (i) they rely on hand-crafted pseudo-sequence extraction, (ii) they do not generalise well to different datasets, which limits the practicality of these approaches. In this paper, we present PerceiverpMHC that is able to learn accurate representations on full-sequences by leveraging efficient transformer based architectures. Additionally, we propose IMGT/RobustpMHC that harnesses the potential of unlabeled data in improving the robustness of peptide-MHC binding predictions through a self-supervised learning strategy. We extensively evaluate RobustpMHC on 8 different datasets and showcase the improvements over the state-of-the-art approaches. Finally, we compile CrystalIMGT, a crystallography verified dataset that presents a challenge to existing approaches due to significantly different peptide-MHC distributions.

bioinformatics↗