bioRxiv · 10.1101/024679
Cookiecutter: a tool for kmer-based read filtering and extraction
Abstract
MotivationKmer-based analysis is a powerful method used in read error correction and implemented in various genome assembly tools. A number of read processing routines include extracting or removing sequence reads from the results of high-throughput sequencing experiments prior to further analysis. Here we present a new approach to sorting or filtering of raw reads based on a provided list of kmers.\n\nResultsWe developed Cookiecutter -- a computational tool for rapid read extraction or removing according to a provided list of k-mers generated from a FASTA file. Cookiecutter is based on the implementation of the Aho-Corasik algorithm and is useful in routine processing of high-throughput sequencing datasets. Cookiecutter can be used for both removing undesirable reads and read extraction from a user-defined region of interest.\n\nAvailabilityThe open-source implementation with user instructions can be obtained from GitHub: https://github.com/ad3002/Cookiecutter.
Explore related subjects
Keep this discovery
Ekaterina Starostina, Gaik Tamazian, Pavel Dobrynin, Stephen O'Brien, Aleksey Komissarov. 2015-08-16. Cookiecutter: a tool for kmer-based read filtering and extraction. https://doi.org/10.1101/024679
Cite the original work for its findings. Save a collection to share your selection of sources.