bioRxiv · 10.1101/2023.02.01.526559
The Master Database of All Possible RNA Sequences and Its Integration with RNAcmap for RNA Homology Search
Abstract
Recent success of AlphaFold2 in protein structure prediction relied heavily on co-evolutionary information derived from homologous protein sequences found in the huge, integrated database of protein sequences (Big Fantastic Database). In contrast, the existing nucleotide databases were not consolidated to facilitate wider and deeper homology search. Here, we built a comprehensive database by including the noncoding RNA sequences from RNAcentral, the transcriptome assembly and metagenome assembly from MG-RAST, the genomic sequences from Genome Warehouse (GWH), and the genomic sequences from MGnify, in addition to NCBIs nucleotide database (nt) and its subsets. The resulting MARS database (Master database of All possible RNA sequences) is 20-fold larger than NCBIs nt database or 60-fold larger than RNAcentral. The new dataset along with a new split-search strategy allows a substantial improvement in homology search over existing state-of-the-art techniques. It also yields more accurate and more sensitive multiple sequence alignments (MSA) than manually curated MSAs from Rfam for the majority of structured RNAs mapped to Rfam. The results indicate that MARS coupled with the fully automatic homology search tool RNAcmap will be useful for improved structural and functional inference of noncoding RNAs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chen, K., Litfin, T., Singh, J., Zhan, J., Zhou, Y.. 2023-02-03. The Master Database of All Possible RNA Sequences and Its Integration with RNAcmap for RNA Homology Search. https://doi.org/10.1101/2023.02.01.526559
Cite the original work for its findings. Save a collection to share your selection of sources.