bioRxiv · 10.64898/2026.06.10.727770
PHI-Reason: evidence-grounded species-level phage-host prediction from structured biological text profiles
Abstract
Biological interaction inference often requires integrating heterogeneous evidence that differs in biological meaning, provenance, specificity, and reliability. Existing computational approaches typically compress such evidence into numerical features, aggregate scores or latent representations before prediction, obscuring the contribution of individual evidence records and limiting interpretation when evidence is conflicting or incomplete. Phage-host interaction prediction provides a demanding test case for this problem, requiring the integration of genomic, functional, and reference-derived evidence with differing coverage, specificity, and reliability. Here we introduce a structured-evidence inference paradigm that retains heterogeneous biological observations as modular, named, and experimentally perturbable records rather than compressing them before prediction. We implement this paradigm in PHI-Reason, which constructs structured evidence profiles from genomic annotations, receptor-binding protein homology, nucleotide-neighbour relationships, alignment-free genomic similarity, and available CRISPR spacer information. A general-purpose large language model (LLM) performs inference directly over these structured profiles without task-specific training or manually engineered evidence-fusion rules. To validate this paradigm, we evaluated it across two distinct biological interaction domains: prokaryotic phage-host prediction and eukaryotic virus-host prediction. Across phage-host benchmarks, PHI-Reason achieved species-level top-1 accuracies of 63.6% and 53.2% on RefSeq-634 and VHDB-3150, respectively, and a multi-host accuracy of 0.571 on the Hi-C cohort, outperforming established numerical methods. Besides, systematic evidence perturbations quantified how individual evidence supported, complemented, or misled inference, identifying nucleotide-neighbour context as the dominant signal. Analyses of LLM intermediate representations using a target-conditioned local Jacobian readout, together with output-grounding analyses, further characterized evidence-dependent inference and quantified the extent to which generated rationales departed from the supplied evidence profiles. Applying the same framework to eukaryotic virus-host prediction using domain-appropriate evidence also preserved high predictive accuracy (66.6%), supporting the generalizability of the proposed paradigm across biological prediction. These results demonstrate that LLMs can serve as an evidence-grounded inference interface, integrating heterogeneous biological evidence while making the boundaries of evidential support directly testable.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhang, Y.-z., Xu, L., Imoto, S.. 2026-06-12. PHI-Reason: evidence-grounded species-level phage-host prediction from structured biological text profiles. https://doi.org/10.64898/2026.06.10.727770
Cite the original work for its findings. Save a collection to share your selection of sources.