Search bioRxiv⌕ Search

Biology subjects

Lorenz, L. J.

Publications and source records attributed to Lorenz, L. J..

3 recordsLinked to original sources

Niche-specific contribution of the branched-chain amino acid biosynthesis protein IlvD in Streptococcus pneumoniae infection

Systemic pneumococcal infections are a major cause of morbidity and mortality. Pneumococci isolated from bacteraemia lack clear genetic signatures distinguishing them from isolates recovered from other disease sites. The bloodstream is an evolutionary dead end, providing no means of onward transmission, and acute bacteraemic events offer limited opportunity for within-host evolution. Nonetheless, we reasoned that it might be possible to identify genetic determinants of virulence in blood through genomic comparison of matched lung and blood isolates from individual infections. Using samples from a mouse invasive pneumonia model, we observed frequent occurrence of low population frequency mutations in ilvD - encoding a branched-chain amino acid (BCAA) biosynthesis protein - in bloodstream isolates. These mutations were absent from the lung-resident bacterial population. The single-nucleotide polymorphisms in ilvD clustered in a short stretch of nucleotides that are highly conserved across bacteria, suggesting they might disrupt protein function. Deletion of ilvD in pneumococcal strain D39 conferred competitive advantages in both in vitro environments mimicking the bloodstream and in mouse and Drosophila systemic infection models. Improved survival of ilvD mutants within phagocytic cells was dependent on iron availability. Conversely, disruption of ilvD reduced fitness in the nasopharynx, suggesting BCAA biosynthesis is required for upper airway colonisation but is a liability in the context of systemic infection. Nasopharyngeal colonisation defects could be rescued by administering intranasal BCAA to infected mice, suggesting loss of capacity for de novo BCAA synthesis in ilvD mutants accounted for reduced upper airway fitness. In support of a critical role for BCAA biosynthesis in the primary commensal lifestyle of pneumococcus, genomic analysis of clinical isolates demonstrated that ilvD is under purifying selection. Collectively, these findings highlight the context-specific role of IlvD and BCAA biosynthesis in infection, with a requirement for protein function in the BCAA-restricted environment of nasopharynx, but fitness benefits deriving from loss of IlvD and its associated iron-sulfur cluster, in the BCAA-rich conditions of blood.

microbiology↗

Rapid and Consistent Genome Clustering for Navigating Bacterial Diversity with Millions of MAGs and Isolates

Bacterial genome and metagenome databases collectively contain over 5 million high-quality assemblies. However, the redundancy of these databases and the limited scalability of existing tools create bottlenecks for fully comprehensive, tree-of-life-scale genomic analyses. A fundamental task is to first break this data into smaller chunks, guided by their genome similarity. However, alignment-based comparative methods struggle to handle more than a few tens of thousands of genomes at a time, making the global organisation computationally complex and expensive. Here, we present gemsparcl (https://github.com/johannahelene/gemsparcl), a tool that clusters bacterial genomes into genomic cohesive units (GCUs), at approximately species-level resolution, over 500 times faster than existing methods. As part of developing gemsparcl, we developed sketchlib.rust, a one-permutation MinHash approach that implements an auxiliary inverted index to further accelerate all-versus-all comparisons. We added a statistical correction for incomplete metagenome-assembled genomes (MAGs) to enable accurate distance estimation and network-based edge quality filtering. After genome completeness quality control, we clustered 5.6 million high-quality bacterial genomes (2.88 million isolates and 2.77 million MAGs) into 92,954 GCUs in [~]14 hours using 48 CPU threads and less than 16.5 GB of memory. Using taxonomic validation of the GCUs, the method achieves very high (99.76%) cluster purity (meaning only one species label occurs per GCU). We demonstrate that the clustering also highlights cases where taxonomic naming can be potentially harmonised or improved. Furthermore, we identify the most frequently reconstructed MAGs that lack a corresponding isolate genome and are thus priorities for culturing. The enhanced speed of gemsparcl enables routine database updates to incorporate the latest genomes. It also makes reference-free microbiome analysis across millions of genomes computationally tractable for the first time.

bioinformatics↗

A reusable model of pangenome selection informs optimal surveillance strategies over vaccine introductions

BackgroundThe human pathogen Streptococcus pneumoniae is a major cause of disease, including pneumonia and meningitis. The introduction of Pneumococcal Conjugate Vaccines (PCVs) initially reduced the burden of disease through a reduction of colonisation by vaccine-targeted serotypes. However, since PCVs only target a proportion of pneumococcal serotypes, they shift intraspecific competition, eventually allowing non-targeted types to replace vaccine types. Understanding the host and pathogen factors causing replacement is important for future vaccine development. Mechanistic understanding of vaccine replacement dynamics is crucial for forecasting and optimisation of genomic surveillance strategies to evaluate realised vaccine effectiveness. MethodsWe developed a mathematical model of the genomic and demographic factors which explain vaccine replacement, used this model to replicate serotype-frequency changes, and investigated cost-effective genomic surveillance strategies. We extended a forward-time model based on the Wright-Fisher model, developing a user-friendly model framework that describes the post-vaccine dynamics of S. pneumoniae populations. Our model describes vaccine replacement as a function of vaccine impact, immigration of new strains, and negative frequency-dependent selection (NFDS) on the accessory genome content. ResultsWe used our model to study vaccine replacement in newly sequenced genomic surveillance data from Nepal, and existing data from the US, and the UK, with distinct surveillance strategies. We showed that the model with NFDS better replicates replacement dynamics than a null model without NFDS, and that NFDS likely only acts on part of the S. pneumoniae accessory genome. We found consistent estimates for vaccination effectiveness across the different study locations and country-specific genes under NFDS, highlighting the importance of conducting genomic surveillance in each country of interest. By simulating data from the model, we showed that an optimal surveillance strategy prioritises per-sampling sample size over sampling frequency for small sampling budgets. ConclusionsOur model can be used to predict vaccine replacement dynamics after PCV introduction, and can be easily reapplied to analyse new data from vaccine introductions or new regions. Our model is available in the R package STUBENTIGER (Studying Balancing Evolution (NFDS) To Investigate Genome Replacement) on GitHub https://github.com/bacpop/Stubentiger.

genomics↗