bioRxiv · 10.1101/2022.06.08.495248
Genomic Benchmarks: A Collection of Datasets for Genomic Sequence Classification
Abstract
In this paper, we propose a collection of curated and easily accessible sequence classification datasets in the field of genomics. The proposed collection is based on a combination of novel datasets constructed from the mining of publicly available databases and existing datasets obtained from published articles. The main aim of this effort is to create a repository for shared datasets that will make machine learning for genomics more comparable and reproducible while reducing the over-head of researchers that want to enter the field. The collection currently contains eight datasets that focus on regulatory elements (promoters, enhancers, open chromatin region) from three model organisms: human, mouse, and roundworm. A simple convolution neural network is also included in a repository and can be used as a baseline model. Benchmarks and the baseline model are distributed as the Python package genomic-benchmarks, and the code is available at https://github.com/ML-Bioinfo-CEITEC/genomic_benchmarks.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gresova, K., Martinek, V., Cechak, D., Simecek, P., Alexiou, P.. 2022-06-10. Genomic Benchmarks: A Collection of Datasets for Genomic Sequence Classification. https://doi.org/10.1101/2022.06.08.495248
Cite the original work for its findings. Save a collection to share your selection of sources.