bioRxiv · 10.64898/2026.02.09.704753
Viral haplotype reconstruction from long reads with virCHap
Abstract
Resolving genomes at the haplotype level for viral populations is crucial for understanding the prevalence of viral diseases and for the development of effective therapeutic treatments. However, viral haplotype reconstruction still presents challenges, such as an unknown number of strains, high inter-strain similarity, repetitive regions, and difficulties in abundance estimation. Here, we developed virCHap, a new reference-based haplotype phasing algorithm for viruses, which applies graph partitioning followed by iteratively quantifiable cluster merging on long-read sequencing data. Benchmarking on simulated and real datasets demonstrates that virCHap outperforms current tools in terms of recall, accurate abundance estimates and read clustering accuracy. On the simulated large-genome VZV experiment, virCHap had a 96.5% recall, 14% higher than the second-best method, and had the most accurate abundance estimates. On a real 5-strain PVY dataset, virCHap had a precision exceeding 92.9%, a recall of over 97%, and a read clustering accuracy of 82%, outperforming the second-best method by 32%. On a real 6-strain SARS-CoV-2 dataset, virCHap achieved >96.9% accuracy, and the most accurate abundance estimates within the spike gene.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gao, Y., Yu, T., Liu, B., Li, G.. 2026-02-10. Viral haplotype reconstruction from long reads with virCHap. https://doi.org/10.64898/2026.02.09.704753
Cite the original work for its findings. Save a collection to share your selection of sources.