Motif selection enables efficient sequence-based classification of non-coding RNA
MotivationNon-coding RNAs (ncRNAs) classification is important for genome annotation and to perform functional analyses of these biological molecules. Efficient methods for large-scale RNA classification remain challenging. Existing methods often rely on structure similarity, and they require huge computing times because they use secondary structure information, which impedes any large-scale use. ResultsWe present a sequence-based method that relies on the computation and the selection of common sequence motifs to provide a set of features for effectively classifying ncRNAs families by a supervised learning approach. The results show that our method achieves an equal or higher accuracy than existing structure-based methods and drastically reduces required computing times. Results also demonstrate that, thanks to an appropriate selection of local sequence motifs, an efficient sequence-based ncRNA classification can be achieved through supervised learning. AvailabilityCode and datasets are available at https://github.com/chegrane/LSC-ncRNA ContactAida.Ouangraoua@usherbrooke.ca Supplementary informationSupplementary data are available at Bioinformatics online.