Search bioRxiv⌕ Search

Biology subjects

Noonan, A. J. C.

Publications and source records attributed to Noonan, A. J. C..

7 recordsLinked to original sources

A comprehensive phage-bacteria interaction atlas links phage lineage and capsule serotype to genome-guided machine learning prediction in Klebsiella pneumoniae

Klebsiella pneumoniae is a WHO critical-priority pathogen for which strain-specific bacteriophages are being explored as precision antimicrobials, yet rapid phage-host matching remains a major barrier to therapeutic deployment. We constructed a comprehensive interaction atlas comprising 84 taxonomically diverse phages and 101 globally sourced, clinically representative K. pneumoniae strains, including multidrug-resistant isolates. Systematic pairwise profiling produced 8,484 interaction measurements, of which 2,656 (31.3%) scored positive for bacterial clearance. Genus was the dominant phage-side determinant of host range, while capsule K-serotype was the strongest host-side determinant of susceptibility; aggregate defense, prophage, plasmid, and antimicrobial-resistance features contributed comparatively little. A genome-guided machine learning model predicted interactions without curated host annotations (AUROC, 0.882; AUPR, 0.765), outperforming a model based only on phage genus and K-serotype and modestly exceeding a curated genomic baseline. The model recovered capsule- and lipopolysaccharide-biosynthesis genes, canonical receptors and defense-associated features as major predictors using SHAP analysis. Feasibility tests of expert- and model-selected cocktails exposed a translational constraint. Although all formulations suppressed growth in vitro, only the specific cocktail whose phages replicated robustly within the murine gut reduced colonization, suggesting in vivo amplification rather than predicted host range as the limiting factor for therapeutic efficacy. Together with the activity of a model-selected cocktail built for an isolate completely excluded from training, these results provide a species-wide resource for K. pneumoniae phage matching and support a hybrid workflow combining genome-based ranking with targeted phenotypic validation.

microbiology↗

ASPIRE: the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems

Microbial communities inhabiting the respiratory tract contribute to health status through interactions with host physiology, immune function, and local environmental conditions. Advances in small subunit ribosomal RNA (SSU or 16S rRNA) gene amplicon sequencing enable culture-independent profiling of microbial communities as amplicon sequence variants (ASVs), revealing links between microbial dysbiosis and respiratory diseases, and the use of mass spectrometry to measure volatile organic compounds (VOCs) in exhaled breath shows emerging promise for biomarker discovery. Here we present ASPIRE, the Amplicon Sequencing Profiler for Investigating Respiratory Ecosystems, an accessible Nextflow workflow for processing, analyzing, and interpreting linked ASV-VOC data from respiratory microbiome studies. ASPIRE is designed to support scalable comparative analysis across respiratory sample types while preserving intermediate file outputs for inspection and reuse within a standardized file structure.

bioinformatics↗

Methodological Impacts on Microbiome Structure and Indicator Status in the Human Lower Respiratory Tract

RationaleLung cancer is the leading cause of cancer-related death globally, and rising incidence among traditionally low-risk individuals intensifies the need for improved early-detection methods that the lower-airway microbiome may inform. ObjectivesTo evaluate how respiratory tract sampling method shapes inferred microbiome structure, and whether bronchial brushing recovers a microbial community ecologically distinct from BAL and oral rinse. MethodsProspective cohort of 33 participants (8 lung cancer, 25 non-cancer controls) underwent oral rinse, bilateral bronchoalveolar lavage (BAL), and bronchial brushing. Microbiome structure was characterized by small subunit ribosomal rRNA gene amplicon sequence variant (ASV) profiling, with indicator-species analysis and SPIEC-EASI correlation-network mapping used to identify ASVs associated with sample type or cancer status. Measurements and Main ResultsSampling method was the dominant axis of variation. BAL communities closely resembled oral rinse (65 jointly indicative ASVs; none shared with brushing), whereas bronchial brushing yielded a distinct but low-biomass signal. After excluding host-derived sequences, two brush-specific indicator ASVs affiliated with Sphingomonadaceae and an uncultured Steroidobacteraceae were identified that co-localized within a single co-occurrence module. Cancer-status effects were not detectable in this pilot cohort, consistent with limited statistical power. ConclusionsSampling method is the primary determinant of inferred respiratory microbiome structure. Bronchial brushing recovers a distinct but low-biomass signal that is obscured when BAL is used in isolation. However, overlap between this low-biomass signal with host and contaminant sequences indicates that confidently resolving a discrete lower-airway community will require deeper sequencing and dedicated contamination controls. These methodological findings directly inform the design of future cancer-and disease-association studies.

microbiology↗

Surface architecture of the bacterial envelope determines phage adsorption route in pathogenic Escherichia coli O157:H7

The outermost surface layers of Gram-negative bacteria determine phage access to terminal receptors, yet their genetic basis has been mapped almost exclusively in laboratory strains that lack them. Here we apply genome-wide RB-TnSeq fitness profiling to four Escherichia coli O157:H7 strains from distinct phylogenetic clades sharing the O157 O-antigen, using 38 phages with terminal receptors previously mapped in E. coli K-12 strain. RB-TnSeq fitness landscapes across all four pathogenic backgrounds were mostly similar, and dominated by surface-associated loci, including the gfc-etk group 4 capsule operon, O-antigen biosynthesis genes, LPS core assembly genes and outer membrane proteins. Disruption of gfc-etk abolished infection in 11 genetically diverse myoviruses, establishing the O-antigen capsule as a widespread required primary recognition substrate. O-antigen loci generated two classes of fitness score patterns. For 10 phages, disruption increased infectivity, indicating it is a barrier to receptor access; for 3 others, disruption abolished infectivity, demonstrating it can also be a primary recognition substrate. Outer membrane protein receptor identity was conserved across laboratory and pathogenic backgrounds, with the same proteins recognized in both K-12 and O157:H7, while glycan layer state determines whether these receptors are reached. These results demonstrate that outer surface glycan layers can act as primary and optional recognition substrates for phage infection, or as physical barriers preventing terminal receptor access. Extending the ability to probe phage-targeted receptors beyond outer membrane proteins provides a framework for incorporating glycan layer state into predictive models of phage-host interactions.

microbiology↗

Comprehensive interaction profiling and machine learning prediction of bacteriophage infectivity across clinically diverse Pseudomonas aeruginosa

The rise of antibiotic-resistant bacterial infections has driven renewed interest in bacteriophage therapy, where viruses that specifically kill bacteria are used as targeted antimicrobials. Pseudomonas aeruginosa, a WHO critical-priority pathogen that causes severe infections in hospitalized and immunocompromised patients, presents a major challenge for phage therapy because of its extraordinary genetic diversity. Phages effective against one bacterial strain often fail against others, and existing cross-resistance-profiling approaches require iterative empirical testing of each new patient isolate. To establish a genome-based framework for rapid phage-isolate matching, we assembled a collection of 95 genomically diverse P. aeruginosa phages representing 20 genera and tested each against 99 genetically diverse clinical isolates, generating 9,405 infection outcome measurements. Bacterial O-antigen serotype emerged as the dominant determinant of strain susceptibility, while defense systems, anti-defense systems, and prophage burden contributed smaller strain-specific effects. The full curated multivariate model explained 47% of strain-susceptibility variance. Machine-learning models integrating these features and pangenome-derived gene clusters reached a per-strain AUROC of 0.86. In an in vivo proof-of-concept test against a single held-out strain, the ML-designed cocktail produced a [~]12-fold greater median CFU reduction than the expert-designed cocktail (q = 0.045), with both cocktails substantially reducing burden relative to the untreated control ([~]113-fold for ML, [~]9-fold for CG; both q < 10{square}3). SHAP analysis of the model identified bacterial surface-architecture genes (LPS biosynthesis, outer membrane proteins, type IV pili) as the dominant predictors, with defense-system content modulating which specific phages succeed against a strain rather than uniformly damping susceptibility. Together, these results establish a genome-based framework for predicting phage susceptibility in genetically diverse clinical isolates.

microbiology↗

Enabling the prediction of phage receptor specificity from genome data

Predicting which receptor a phage binds to from genome sequence alone has remained an intractable challenge, principally because the experimental phenotypic data required to train and validate predictive models have not been available at sufficient scale. Here we address this by conducting 1,050 genome-wide genetic screens across 255 taxonomically diverse Escherichia coli dsDNA phages, assigning host receptors to 193 phages across 19 receptor classes. Comparative genomics and AlphaFold3 structural modelling resolved the sequence determinants of specificity to defined receptor-binding protein domains and individual residues. Machine learning models trained on this dataset predicted host receptor identity from phage genome sequence alone without prior annotation of receptor-binding genes, achieving perfect precision and greater than 80% recall on 49 independently validated phages, and yielding predictions for 1,060 of 1,875 E. coli phage genomes in NCBI. Domain swaps redirected receptor specificity as predicted, and a single amino acid substitution proved both necessary and sufficient to switch recognition between two distinct porins. These results demonstrate that systematic phenotyping at scale makes sequence-based prediction of molecular interaction specificity tractable, with direct implications for phage-based medicine, microbiome engineering and the broader challenge of inferring host-pathogen interaction outcomes from sequence.

microbiology↗

Phylogeny-agnostic strain-level prediction of phage-host interactions from genomes

Bacteriophages offer promising alternatives to antibiotics for treating drug-resistant infections and engineering microbiomes, but applications are limited by inability to select phages infecting specific bacterial strains. Selecting suitable phages requires either one-to-one experimental assays or strain-level predictions of phage-host interactions. Existing computational approaches either predict host taxonomy at broad ranks unsuitable for strain-level targeting or require species-specific mechanistic knowledge limiting generalizability. Here, we present a phylogenyagnostic machine learning framework predicting strain-level phage-host interactions across diverse bacterial genera from genome sequences alone. Systematically optimizing the workflow over 13.2 million training runs across six datasets (115,037 interactions, 949 bacterial strains, 518 phages), we achieved performance matching species-specific methods (AUROC 0.67-0.94) while eliminating phylogenetic constraints. Comprehensive feature engineering identifies biologically interpretable genetic determinants while minimizing overfitting in sparse, imbalanced datasets. Experimental validation through 1,240 novel interactions confirmed generalizability (AUROC 0.84), while genome-wide RB-TnSeq screens verified that 68.6% of experimentally identified infection mediators were captured computationally, including receptors and cell wall biosynthesis pathways. Model-guided cocktail design achieved up to 97.5% bacterial coverage with five phages, and up to a 3.1-fold improvement in single-phage selection over promiscuity-based selection. This platform enables rational phage therapy design and precision microbiome engineering with applications in combating antimicrobial resistance across clinical, agricultural, and industrial contexts.

microbiology↗