Search bioRxiv⌕ Search

Biology subjects

Maucourt, F.

Publications and source records attributed to Maucourt, F..

3 recordsLinked to original sources

A comprehensive phage-bacteria interaction atlas links phage lineage and capsule serotype to genome-guided machine learning prediction in Klebsiella pneumoniae

Klebsiella pneumoniae is a WHO critical-priority pathogen for which strain-specific bacteriophages are being explored as precision antimicrobials, yet rapid phage-host matching remains a major barrier to therapeutic deployment. We constructed a comprehensive interaction atlas comprising 84 taxonomically diverse phages and 101 globally sourced, clinically representative K. pneumoniae strains, including multidrug-resistant isolates. Systematic pairwise profiling produced 8,484 interaction measurements, of which 2,656 (31.3%) scored positive for bacterial clearance. Genus was the dominant phage-side determinant of host range, while capsule K-serotype was the strongest host-side determinant of susceptibility; aggregate defense, prophage, plasmid, and antimicrobial-resistance features contributed comparatively little. A genome-guided machine learning model predicted interactions without curated host annotations (AUROC, 0.882; AUPR, 0.765), outperforming a model based only on phage genus and K-serotype and modestly exceeding a curated genomic baseline. The model recovered capsule- and lipopolysaccharide-biosynthesis genes, canonical receptors and defense-associated features as major predictors using SHAP analysis. Feasibility tests of expert- and model-selected cocktails exposed a translational constraint. Although all formulations suppressed growth in vitro, only the specific cocktail whose phages replicated robustly within the murine gut reduced colonization, suggesting in vivo amplification rather than predicted host range as the limiting factor for therapeutic efficacy. Together with the activity of a model-selected cocktail built for an isolate completely excluded from training, these results provide a species-wide resource for K. pneumoniae phage matching and support a hybrid workflow combining genome-based ranking with targeted phenotypic validation.

microbiology↗

Comprehensive interaction profiling and machine learning prediction of bacteriophage infectivity across clinically diverse Pseudomonas aeruginosa

The rise of antibiotic-resistant bacterial infections has driven renewed interest in bacteriophage therapy, where viruses that specifically kill bacteria are used as targeted antimicrobials. Pseudomonas aeruginosa, a WHO critical-priority pathogen that causes severe infections in hospitalized and immunocompromised patients, presents a major challenge for phage therapy because of its extraordinary genetic diversity. Phages effective against one bacterial strain often fail against others, and existing cross-resistance-profiling approaches require iterative empirical testing of each new patient isolate. To establish a genome-based framework for rapid phage-isolate matching, we assembled a collection of 95 genomically diverse P. aeruginosa phages representing 20 genera and tested each against 99 genetically diverse clinical isolates, generating 9,405 infection outcome measurements. Bacterial O-antigen serotype emerged as the dominant determinant of strain susceptibility, while defense systems, anti-defense systems, and prophage burden contributed smaller strain-specific effects. The full curated multivariate model explained 47% of strain-susceptibility variance. Machine-learning models integrating these features and pangenome-derived gene clusters reached a per-strain AUROC of 0.86. In an in vivo proof-of-concept test against a single held-out strain, the ML-designed cocktail produced a [~]12-fold greater median CFU reduction than the expert-designed cocktail (q = 0.045), with both cocktails substantially reducing burden relative to the untreated control ([~]113-fold for ML, [~]9-fold for CG; both q < 10{square}3). SHAP analysis of the model identified bacterial surface-architecture genes (LPS biosynthesis, outer membrane proteins, type IV pili) as the dominant predictors, with defense-system content modulating which specific phages succeed against a strain rather than uniformly damping susceptibility. Together, these results establish a genome-based framework for predicting phage susceptibility in genetically diverse clinical isolates.

microbiology↗

Enabling the prediction of phage receptor specificity from genome data

Predicting which receptor a phage binds to from genome sequence alone has remained an intractable challenge, principally because the experimental phenotypic data required to train and validate predictive models have not been available at sufficient scale. Here we address this by conducting 1,050 genome-wide genetic screens across 255 taxonomically diverse Escherichia coli dsDNA phages, assigning host receptors to 193 phages across 19 receptor classes. Comparative genomics and AlphaFold3 structural modelling resolved the sequence determinants of specificity to defined receptor-binding protein domains and individual residues. Machine learning models trained on this dataset predicted host receptor identity from phage genome sequence alone without prior annotation of receptor-binding genes, achieving perfect precision and greater than 80% recall on 49 independently validated phages, and yielding predictions for 1,060 of 1,875 E. coli phage genomes in NCBI. Domain swaps redirected receptor specificity as predicted, and a single amino acid substitution proved both necessary and sufficient to switch recognition between two distinct porins. These results demonstrate that systematic phenotyping at scale makes sequence-based prediction of molecular interaction specificity tractable, with direct implications for phage-based medicine, microbiome engineering and the broader challenge of inferring host-pathogen interaction outcomes from sequence.

microbiology↗