bioRxiv · 10.1101/420745
A new statistic for efficient detection of repetitive sequences
Abstract
Detecting sequences containing repetitive regions is a basic bioinformatics task with many applications. Several methods have been developed for various types of repeat detection tasks. An efficient generic method for detecting all types of repetitive sequences is still desirable.\n\nInspired by the excellent properties and successful applications of the D2 family of statistics in comparative analyses of genomic sequences, we developed a new statistic [Formula] that can efficiently discriminate sequences with or without repetitive regions. Using the statistic, we developed an algorithm of linear complexity in both computation time and memory usage for detecting all types of repetitive sequences in multiple scenarios, including finding candidate CRISPR regions from bacterial genomic or metagenomics sequences. Simulation and real data experiments showed that the method works well on both assembled sequences and unassembled short reads.
Source connections
Explore related subjects
Keep this discovery
Chen, S., Sun, F., Waterman, M. S., Zhang, X.. 2018-09-18. A new statistic for efficient detection of repetitive sequences. https://doi.org/10.1101/420745
Cite the original work for its findings. Save a collection to share your selection of sources.