bioRxiv · 10.1101/2023.07.03.547471
PanTA: An ultra-fast method for constructing large and growing microbial pangenomes
Abstract
Pangenome analysis is an indispensable step in bacterial genomics to address the high variability of bacteria genomes. However, speed and scalability remain a challenge for pangenome inference software tools to cope with the fast-growing genomic collections. We present PanTA, a software package for constructing the pangenomes of large bacterial collections. We show that PanTA exhibits an unprecedented multiple times more efficient than the current state-of-the-arts while maintaining a similar pangenome accuracy. In addition, PanTA introduces a novel mechanism to construct the pangenome progressively where new samples are added into an existing pangenome without rebuilding the accumulated collection from scratch. In the progressive mode, PanTA is demonstrated to consume orders of magnitude less computational resource than existing solutions in managing the pangenomes of growing microbial datasets. We further show that PanTA can build the pangenome of the entire collection of >28000 Escherichia coli genomes from the RefSeq database on a laptop computer in 32 hours, highlighting the scalability and practicality of PanTA.The software is open source and is publicly available at https://github.com/amromics/panta under an MIT license.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Le, D. Q., Nguyen, T. A., Nguyen, T. T., Do, V. H., Nguyen, C. H., Phung, H. T., Ho, T. H., Vo, N. S., Nguyen, T., Nguyen, H. A., Cao, M. D.. 2023-07-03. PanTA: An ultra-fast method for constructing large and growing microbial pangenomes. https://doi.org/10.1101/2023.07.03.547471
Cite the original work for its findings. Save a collection to share your selection of sources.