bioRxiv · 10.1101/2020.12.21.423849
Increased accuracy and speed in whole genome bisulfite read mapping using a two-letter alphabet
Abstract
DNA cytosine methylation is an important epigenomic mark with a wide range of functions across many organisms. Whole genome bisulfite sequencing (WGBS) is the gold standard to interrogate cyto-sine methylation genome-wide. Algorithms used to map WGBS reads often encode the four-base DNA alphabet with three letters by reducing two bases to a common letter. This encoding substantially reduces the entropy of nucleotide frequencies in the resulting reference genome. Within the paradigm of read mapping by first filtering possible candidate alignments, reduced entropy of the reference can increase the required computing effort. We introduce another bisulfite mapping algorithm (abismal), based on the idea of encoding a four-letter DNA sequence as only two letters, one for purines and one for pyrimidines. We show that this encoding has greater specificity when subsequences are selected from reads for filtration. Through the two-letter encoding, the abismal software tool maps reads in less time and using less memory than most WGBS read mapping software tools, while attaining similar accuracy. This allows in silico methylation analysis to be performed in a wider range of computing machines with limited hardware settings.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
de Sena Brandine, G., Smith, A. D.. 2020-12-22. Increased accuracy and speed in whole genome bisulfite read mapping using a two-letter alphabet. https://doi.org/10.1101/2020.12.21.423849
Cite the original work for its findings. Save a collection to share your selection of sources.