Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.02.10.705007

A stable, hierarchical LIN code system for Campylobacter jejuni and Campylobacter coli: A unified genomic nomenclature for lineage-level typing and global surveillance.

Abstract

Campylobacter remains the leading cause of bacterial gastroenteritis worldwide, with C. jejuni accounting for around 90% of infection and C. coli accounting for most of the rest. Seven-locus multilocus sequence typing (MLST) has improved our understanding of host association and population structure, whilst core genome MLST (cgMLST), enables investigation of transmission events at high-resolution. However, the lack of a stable and standardised nomenclature for clustering of cgMLST data has limited reproducibility and long-term comparability between studies. Here we introduce a joint, hierarchical Life Identification Number (LIN) code system that provides reproducible, multi-level genomic identifiers for C. jejuni and C. coli lineages. Using an updated cgMLST v2 scheme (1,142 loci) and globally representative datasets of high-quality genomes selected from over 53,000 assemblies in the Campylobacter PubMLST database (https://pubmlst.org/organisms/campylobacter-jejunicoli), we firstly defined LIN codes on a dataset of 5,664 genomes. Pairwise allelic distances were computed using MSTclust, and 18 nested thresholds were defined through silhouette, adjusted Wallace and adjusted Rand Index (ARI) statistics to capture the population structure from species to outbreak level resolution. The LIN thresholds were then validated using a second dataset of 1,781 genomes from PubMLST and applied to a large water-associated outbreak dataset from New Zealand in 2016, containing clinical and ecological genomes. Further application of LIN codes was demonstrated by analyses of the C. jejuni ST-21 clonal complex and ST-6175 isolates, as well as the broader population structure of C. coli, using data from PubMLST. Across all datasets, LIN clusters were stable, largely monophyletic, and back-compatible with existing nomenclature, accurately distinguishing host-adapted and outbreak-associated lineages. By embedding cgMLST data within a stable and scalable nomenclature, the Campylobacter LIN system delivers consistent, automated genome-to-lineage assignment. This unified framework bridges population genetics and applied surveillance, enabling robust, real-time comparison of Campylobacter isolates across sources, studies, and time. Impact statementHuman cases of Campylobacter worldwide continue unabated. Tracing the source of Campylobacter infection is particularly challenging given the sporadic or multi-source nature of outbreaks, with potential transmission from foodborne, animal or environmental sources. Seven-locus MLST has greatly improved our broad understanding of Campylobacter population structure. However, whilst high-resolution cgMLST alleles and STs themselves do not change, longitudinal cluster analyses of cgMLST data have lacked a stable nomenclature, rendering them unsuitable for robust and comparable surveillance over time. Life Identification Number (LIN) codes provide a solution to this problem, establishing an automated and scalable nomenclature derived directly from cgMLST profiles, that is stable over time. We have implemented a joint C. jejuni and C. coli LIN code scheme in PubMLST, with scripts for real-time lineage assignment. LIN codes are back-compatible with existing MLST nomenclature, and we demonstrate their added practical value for exploring population structure and high-resolution outbreak investigation. LIN codes support surveillance of Campylobacter in a One Health context, by enabling consistent typing at multiple levels across different sources, laboratories and time. Data summary1. The isolate collections used to develop the LIN codes are publicly available and searchable as individual projects on the PubMLST database (https://pubmlst.org). O_LILIN code development (Dataset 1) (n=5,664 isolates, up to 200 isolates per clonal complex) C_LIO_LILIN code validation (Dataset 2) (n=1,781 isolates), up to 50 isolates per clonal complex C_LIO_LIOutbreak investigation (Dataset 3): New Zealand 2016 Havelock North waterborne outbreak, Gilpin et al (n=161 isolates) [1] C_LIO_LIPopulation structure exploration (clonal complex) (Dataset 4): ST-21 complex (n=1800 isolates, up to 100 isolates randomly selected from each country) C_LIO_LIPopulation structure exploration (sequence type) (Dataset 5): ST-6175 (n=321 isolates, genomes with good cgMLST v2 annotation) C_LI 2. The software for LIN code development is publicly available as follows: O_LIMSTclust for pairwise distance matrices https://gitlab.pasteur.fr/GIPhy/MSTclust [2] C_LIO_LIPython script to define LIN codes in a local dataset; (https://gitlab.pasteur.fr/BEBP/LINcoding) C_LIO_LIBIGSdb Perl script to define LIN codes from cgMLST profiles on the PubMLST database; (https://github.com/kjolley/BIGSdb/blob/develop/scripts/maintenance/lincodes.pl) C_LI

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Parfitt, K. M., Pascoe, B., Jolley, K. A., Douglas, A., Goforth, M. P., Sheppard, S. K., Maiden, M. C. J., Colles, F. M.. 2026-02-10. A stable, hierarchical LIN code system for Campylobacter jejuni and Campylobacter coli: A unified genomic nomenclature for lineage-level typing and global surveillance.. https://doi.org/10.64898/2026.02.10.705007

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Matrix-controlled emergence of biofilm architecture shapes antimicrobial survival

Biofilms are structured microbial communities whose extracellular matrix is widely regarded as a basis of their protection against antimicrobial compounds. Yet how matrix production by individual bacteria gives rise to collective architecture and antimicrobial protection remains poorly understood. Here, we systematically varied expression of the master biofilm regulator csgD in Salmonella enterica and found that increasing matrix production reorganizes biofilms from dense, isotropic packings into sparse, nematically aligned communities by altering cell-cell interactions. By combining experimentally measured biofilm architectures with reaction-diffusion modeling, we show that these structural changes produce distinct patterns of antimicrobial killing, ranging from preferential killing near the liquid-biofilm interface to more uniform killing throughout the community. Consequently, increasing matrix production unexpectedly reduces antimicrobial survival by shifting the biofilm into different transport regimes, while strain-specific physiological differences further modulate antimicrobial depletion. Rather than acting as a passive barrier, EPS therefore shapes antimicrobial susceptibility by reorganizing biofilm architecture and its transport properties. EPS thus provides a physical link between molecular regulation, collective architecture and antimicrobial survival, providing a quantitative framework for understanding how cellular matrix production generates emergent biofilm function.

microbiology↗

Mapping virulence-associated protein interaction networks reveals regulators of thermotolerance in Cryptococcus neoformans

Protein-protein interactions (PPIs) influence critical biological processes in pathogenic microorganisms, such as the human fungal pathogen, Cryptococcus neoformans. Fungal thermotolerance and stress response pathways are key virulence determinants that directly impact pathogen adaptation and survival and the infection process. To establish a comprehensive baseline of PPIs in C. neoformans and explore these interactions to infer functional roles for uncharacterized proteins, we applied size exclusion chromatography coupled with mass spectrometry to the secreted and cellular proteomes of the fungi. As a result, 216 and 1699 unique proteins were identified across 24 secretome and proteome fractions, respectively. The predicted secretome networks included expected proteins associated with vesicles and virulence, indicating a role in extracellular defense. Whereas the cryptococcal proteome highlighted interactions among proteins with defined roles in fungal virulence for protein stability and thermotolerance, including two previously uncharacterized proteins, CNAG_00287 and CNAG_05199, putatively involved in complex formation with heat-shock proteins (HSP). Based on sequence and structure homology, we propose that CNAG_00287 is a tetratricopeptide repeat-containing co-chaperone that modulates Hsp 70 activity and CNAG_05199 functions as a Hsp70. We validated the thermotolerance role of CNAG_00287 in heat-related stress, as its absence significantly impaired fungal growth in nutrient-limited media at 37 {degrees}C. Together, this work resolves virulence-associated PPIs within C. neoformans and reveals new molecular regulators of thermotolerance that underpin fungal pathogenicity.

microbiology↗

Environmental filtering and host identity collectively shape root-associated microbiomes of Ericaceae and ectomycorrhizal plants in fumarole fields

Background Symbiosis with microbes is a key strategy that has enabled plants to colonize extreme environments. Since the benefits conferred by root-associated microbes depend on both environmental conditions and host-microbe combinations, plant adaptation to harsh environments is closely linked to the assembly of root microbial communities. Understanding how environmental and host filtering jointly shape these communities is therefore fundamental to elucidating the mechanisms underlying plant adaptation to extreme environments. Results In this study, we investigated the differentiation of root-associated prokaryotic and fungal communities and individual operational taxonomic units (OTUs) across two contrasting habitats surrounding fumaroles, solfatara-field and forest-edge habitats, and six dominant Ericaceae and ectomycorrhizal plant taxa. Prokaryotic and fungal OTUs rarely exhibited strong preferences for both habitat and host identity. Instead, many of prokaryotic and fungal OTUs specialized to one of these niches, collectively generating root microbial communities differentiated by both factors. Nonetheless, striking specializations in habitat and host niches were observed in the fungal family Hyaloscyphaceae (Helotiales). To gain insight into the evolutionary basis of microbial specialization, we examined phylogenetic signals in preference phenotypes. The resulting weak phylogenetic signals in these preference phenotypes further suggest that this fungal clade has undergone substantial ecological divergence. Conclusion Overall, our findings indicate that root-associated microbial communities in extreme environments are assembled through the accumulation of microbial taxa specialized to either habitat or host, and that strong ecological specialization in fungi can arise with little phylogenetic constraint.

microbiology↗