bioRxiv · 10.1101/648683
Algorithms for efficiently collapsing reads with Unique Molecular Identifiers
Abstract
BackgroundUnique Molecular Identifiers (UMI) are used in many experiments to find and remove PCR duplicates. Although there are many tools for solving the problem of deduplicating reads based on their finding reads with the same alignment coordinates and UMIs, many tools either cannot handle substitution errors, or require expensive pairwise UMI comparisons that do not efficiently scale to larger datasets.\n\nResultsWe formulate the problem of deduplicating UMIs in a manner that enables optimizations to be made, and more efficient data structures to be used. We implement our data structures and optimizations in a tool called UMICollapse, which is able to deduplicate over one million unique UMIs of length 9 at a single alignment position in around 26 seconds.\n\nConclusionsWe present a new formulation of the UMI deduplication problem, and show that it can be solved faster, with more sophisticated data structures.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Liu, D.. 2019-05-24. Algorithms for efficiently collapsing reads with Unique Molecular Identifiers. https://doi.org/10.1101/648683
Cite the original work for its findings. Save a collection to share your selection of sources.