bioRxiv · 10.1101/2020.05.24.113852
Predicting Alignment Distances via Continuous Sequence Matching
Abstract
Quantifying pairwise sequence similarities is a key step in metagenomics studies. Alignment-free methods provide a computationally efficient alternative to alignment-based methods for large-scale sequence analysis. Several neural network-based methods have recently been developed for this purpose. However, existing methods do not perform well on sequences of varying lengths and are sensitive to the presence of insertions and deletions. In this paper, we describe the development of a new method, referred to as AsMac, that addresses the aforementioned issues. We proposed a novel neural network structure for approximate string matching for the extraction of pertinent information from biological sequences and developed an efficient gradient computation algorithm for training the constructed neural network. We performed a large-scale benchmark study using real-world data that demonstrated the effectiveness and potential utility of the proposed method. The open-source software for the proposed method and trained neural-network models for some commonly used metagenomics marker genes were developed and are freely available at www.acsu.buffalo.edu/~yijunsun/lab/AsMac.html.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chen, J., Yang, L., Li, L., Sun, Y.. 2020-05-27. Predicting Alignment Distances via Continuous Sequence Matching. https://doi.org/10.1101/2020.05.24.113852
Cite the original work for its findings. Save a collection to share your selection of sources.