Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.03.27.645760

Disentangling the impact of abiotic and biotic environmental factors and dispersal dynamics on bacterial pangenome fluidity

Abstract

Understanding how pangenomes originate and evolve is crucial for predicting evolutionary trajectories and uncovering ecological interactions of bacterial pathogens. Pangenome fluidity has been attributed to adaptive evolution, yet the underlying ecological drivers for bacterial pathogens persisting in natural reservoirs remain poorly understood. Listeria monocytogenes (Lm), a foodborne pathogen causing fatal listeriosis, serves as an ideal model for investigating the ecological mechanisms underlying pangenome fluidity in bacterial pathogens due to its high evolutionary divergence, broad ecological versatility, and significant public health concern. Through pangenome analysis of 177 Lm isolates representing three evolutionary lineages (I, II, and III) that we isolated from soils across the United States, we found that substantial genome variation was strongly associated with climatic factors (e.g. precipitation and temperature), soil properties (e.g. aluminum, pH, and molybdenum), and bacterial community composition, particularly Nitrospirae, Planctomycetes, Acidobacteria, and Cyanobacteria. These factors exerted selective pressure across many gene functions, with pronounced effects on genes involved in cell envelope synthesis, defense mechanisms, and replication, recombination, and repair. Among Lm lineages occupying varied habitats, distinct pangenome properties were observed. Lineage III exhibited a highly fluid pangenome, which was attributed to local adaptation to nutrient-limited conditions and strong dispersal limitation. In contrast, lineage I maintained a conserved pangenome, likely due to frequent homogenizing dispersal. Consistent with these dispersal patterns, we identified an elevated risk of soil-to-human transmission in lineage I, evidenced by epidemiological links between three soil-derived and 17 clinical isolates. Collectively, this study reveals the pivotal role of environmental selection imposed by both abiotic factors and bacterial communities in governing the adaptive pangenome evolution in bacterial pathogens. It also highlights significant differences in pangenome flexibility, ecological niches, and transmission dynamics across lineages of the same pathogen species, underscoring the need for tailored source tracking strategies. AUTHOR SUMMARYStudying the full set of genes found in different strains of a bacterium (i.e. pangenome) helps us understand how bacterial pathogens develop and adapt to changes in the environment. Here, we focused on Listeria monocytogenes (Lm), a pathogen capable of spreading through food and surviving in diverse environments, to understand how environmental factors and the way that bacteria move across locations can influence the pangenome content in this important bacterium. By analyzing the genomes of 177 Lm strains representing three evolutionary lineages (I, II, and III) collected from soils across the United States, we found that variation in climate, soil chemistry, and surrounding bacteria (e.g., Nitrospirae) was closely linked to genetic differences among strains. These environmental conditions seemed to affect genes that help build the cell envelop, protect the bacteria from harm, and fix damaged DNA. We also observed different levels of genome flexibility across Lm lineages which were found to be related to how they move across different locations. Lineage III showed evidence of barriers to spreading, which may enhance genetic differentiation across populations, leading to a more flexible pangenome. In contrast, lineage I appeared to spread more readily and was epidemiologically linked to human clinical cases, which may facilitate genetic exchange that reduce pangenome diversity. This study shows that both non-living environmental conditions--like precipitation and pH--and nearby groups of bacteria play a big role in shaping how bacterial pathogens change their genes to survive. It also highlights that different subtypes of the same pathogen can have different gene flexibility and spread in different ways, calling for specific biocontrol measures.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Goh, Y.-X., Hardeep, F., Zhang, H., Liao, J.. 2025-04-01. Disentangling the impact of abiotic and biotic environmental factors and dispersal dynamics on bacterial pangenome fluidity. https://doi.org/10.1101/2025.03.27.645760

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Structural variation in repeat elements is widespread in normal human tissues and in tumorigenesis

Somatic mosaicism contributes to genomic variation, yet postzygotic structural variants remain under-characterized. We performed long- and short-read WGS from multiple individuals (n=47 normal tissues; n=168 samples) and identified mosaic structural variants in all individuals and germ layers, impacting a median 285.2 kb/genome. Nearly half of breakpoints were independently validated, with tissue distributions reflecting both early and late developmental origins. Most mosaic variants were repeat-mediated and 8.3% overlapped functional elements, an enrichment compared to germline variants. To extend these analyses in samples where long-read sequencing is infeasible, we measured repeat alterations from short-read sequencing, recapitulating mosaic tissue-specific differences. We characterized tumor- and tissue- specific variation in repeats across 15 cancer types and found tumor-related repeat variation to be similar in scale to that of normal mosaic variation. Tracking repeat changes in cell-free DNA provided a noninvasive approach for tumor monitoring. Our analyses revealed widespread repeat-driven structural variation in health and disease.

genomics↗

RNA isoform-resolved multiplexed sequencing with bioorthogonal barcoding

RNA isoform dysregulation drives disease pathogenesis and is the target of FDA-approved splice-switching therapeutics. However, multiplexed sequencing methods discard splice junction information because only 3' termini are barcoded and counted. Here, we repurpose acylation and click chemistries to conjugate bioorthogonal barcodes (bobcodes) directly onto multiple internal positions along cellular RNAs. Bobcoded RNAs from multiple samples are pooled for multiplexed cDNA synthesis, during which reverse transcriptase switches from each RNA template onto its tethered bobcode with greater than 99% accuracy in species mixing experiments. Bobcode attachment intervals set cDNA insert sizes without a library fragmentation step, and priming with poly(dT) or random hexamers selects between 3'-end counting and full-length isoform capture. A bioorthogonal barcode-sequencing (BOB-seq v0.1) drug screen identifies transcriptome-wide on- and off-target RNA splicing effects and outperforms existing multiplexing RNA sequencing methods in workflow simplicity, sample-to-sample variability, and barcoding accuracy. Bobcodes add isoform resolution to scalable multiplexed RNA sequencing.

genomics↗

Structural polymorphism and population-variable coding capacity of HERV-K(HML-2) in human pangenomes

Approximately 8% of the human genome is derived from ancient retroviral infections. The most recently integrated of these endogenous retroviruses is the HERV-K(HML-2) clade, whose expression has been associated with cancer, amyotrophic lateral sclerosis, and embryogenesis. Studies of HERV expression, particularly HML-2, have relied predominantly on short-read sequencing. However, the high similarity among HML-2 proviruses prevents many short reads from being assigned uniquely to individual loci. We therefore compared haplotype-resolved long-read genome assemblies from 292 donors to resolve variation in proviral structure and coding capacity. Several loci previously thought to be fixed were structurally polymorphic. Tandem arrays occurred at 13 loci and contained up to six proviral copies in a single array. At 8q11.23, we identified a previously undescribed full-length provirus in one haplotype. All 583 other haplotypes carried a solo-LTR. We found that standard reference genomes failed to represent the coding capacity retained in many individuals, whose proviruses contained intact open reading frames despite disruptive mutations in the reference sequences. Short-read genotypes left 32.5% of the tested donor-variant pairs unresolved at sites associated with viral reading frames. These findings show why HML-2 expression must be interpreted in the context of the structural and coding alleles each individual carries.

genomics↗