Search bioRxiv⌕ Search

Biology subjects

Rivera-Lopez, E. O.

Publications and source records attributed to Rivera-Lopez, E. O..

3 recordsLinked to original sources

Surface architecture of the bacterial envelope determines phage adsorption route in pathogenic Escherichia coli O157:H7

The outermost surface layers of Gram-negative bacteria determine phage access to terminal receptors, yet their genetic basis has been mapped almost exclusively in laboratory strains that lack them. Here we apply genome-wide RB-TnSeq fitness profiling to four Escherichia coli O157:H7 strains from distinct phylogenetic clades sharing the O157 O-antigen, using 38 phages with terminal receptors previously mapped in E. coli K-12 strain. RB-TnSeq fitness landscapes across all four pathogenic backgrounds were mostly similar, and dominated by surface-associated loci, including the gfc-etk group 4 capsule operon, O-antigen biosynthesis genes, LPS core assembly genes and outer membrane proteins. Disruption of gfc-etk abolished infection in 11 genetically diverse myoviruses, establishing the O-antigen capsule as a widespread required primary recognition substrate. O-antigen loci generated two classes of fitness score patterns. For 10 phages, disruption increased infectivity, indicating it is a barrier to receptor access; for 3 others, disruption abolished infectivity, demonstrating it can also be a primary recognition substrate. Outer membrane protein receptor identity was conserved across laboratory and pathogenic backgrounds, with the same proteins recognized in both K-12 and O157:H7, while glycan layer state determines whether these receptors are reached. These results demonstrate that outer surface glycan layers can act as primary and optional recognition substrates for phage infection, or as physical barriers preventing terminal receptor access. Extending the ability to probe phage-targeted receptors beyond outer membrane proteins provides a framework for incorporating glycan layer state into predictive models of phage-host interactions.

microbiology↗

Enabling the prediction of phage receptor specificity from genome data

Predicting which receptor a phage binds to from genome sequence alone has remained an intractable challenge, principally because the experimental phenotypic data required to train and validate predictive models have not been available at sufficient scale. Here we address this by conducting 1,050 genome-wide genetic screens across 255 taxonomically diverse Escherichia coli dsDNA phages, assigning host receptors to 193 phages across 19 receptor classes. Comparative genomics and AlphaFold3 structural modelling resolved the sequence determinants of specificity to defined receptor-binding protein domains and individual residues. Machine learning models trained on this dataset predicted host receptor identity from phage genome sequence alone without prior annotation of receptor-binding genes, achieving perfect precision and greater than 80% recall on 49 independently validated phages, and yielding predictions for 1,060 of 1,875 E. coli phage genomes in NCBI. Domain swaps redirected receptor specificity as predicted, and a single amino acid substitution proved both necessary and sufficient to switch recognition between two distinct porins. These results demonstrate that systematic phenotyping at scale makes sequence-based prediction of molecular interaction specificity tractable, with direct implications for phage-based medicine, microbiome engineering and the broader challenge of inferring host-pathogen interaction outcomes from sequence.

microbiology↗

Phylogeny-agnostic strain-level prediction of phage-host interactions from genomes

Bacteriophages offer promising alternatives to antibiotics for treating drug-resistant infections and engineering microbiomes, but applications are limited by inability to select phages infecting specific bacterial strains. Selecting suitable phages requires either one-to-one experimental assays or strain-level predictions of phage-host interactions. Existing computational approaches either predict host taxonomy at broad ranks unsuitable for strain-level targeting or require species-specific mechanistic knowledge limiting generalizability. Here, we present a phylogenyagnostic machine learning framework predicting strain-level phage-host interactions across diverse bacterial genera from genome sequences alone. Systematically optimizing the workflow over 13.2 million training runs across six datasets (115,037 interactions, 949 bacterial strains, 518 phages), we achieved performance matching species-specific methods (AUROC 0.67-0.94) while eliminating phylogenetic constraints. Comprehensive feature engineering identifies biologically interpretable genetic determinants while minimizing overfitting in sparse, imbalanced datasets. Experimental validation through 1,240 novel interactions confirmed generalizability (AUROC 0.84), while genome-wide RB-TnSeq screens verified that 68.6% of experimentally identified infection mediators were captured computationally, including receptors and cell wall biosynthesis pathways. Model-guided cocktail design achieved up to 97.5% bacterial coverage with five phages, and up to a 3.1-fold improvement in single-phage selection over promiscuity-based selection. This platform enables rational phage therapy design and precision microbiome engineering with applications in combating antimicrobial resistance across clinical, agricultural, and industrial contexts.

microbiology↗