bioRxiv · 10.64898/2026.09.17.752512
Lossless compression of protein databases for efficient and accurate metagenomic sequence classification with Centrifuger
Abstract
We present a lossless compression algorithm for indexing a protein database while supporting fast taxonomic classification in the method Centrifuger. The algorithm is a new scheme of the previously proposed run-block compression algorithm to reduce the size of the Ferragina-Manzini (FM) index, and it scales better with alphabet size than the original. On the RefSeq prokaryotic and viral protein sequences, Centrifuger reduces the memory footprint by over a third compared to the method Kaiju that builds on a plain FM-index, while having comparable running time. Furthermore, the compressed FM-index is lossless and can locate matches of arbitrary length, which helps Centrifuger achieve greater accuracy than Kraken2, a k-mer-based taxonomic classification method. We leverage the computational efficiency of Centrifuger to create an index of size 182 GB for classifying the reads against the full nr database that contains about 250 billion amino acid characters. Using this index, Centrifuger reveals different SARS-CoV-2 infection states and viral transcriptome profiles across human cell types from single-cell RNA-seq data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Song, L.. 2026-09-24. Lossless compression of protein databases for efficient and accurate metagenomic sequence classification with Centrifuger. https://doi.org/10.64898/2026.09.17.752512
Cite the original work for its findings. Save a collection to share your selection of sources.