bioRxiv · 10.1101/2023.07.19.549462
Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery
Abstract
DNA variation analysis has become indispensable in many aspects of modern biomedicine, most prominently in the comparison of normal and tumor samples. Thousands of samples are collected in local sequencing efforts and public databases requiring highly scalable, portable, and automated workflows for streamlined processing. Here, we present nf-core/sarek 3, a well-established, comprehensive variant calling and annotation pipeline for germline and somatic samples. It is suitable for any genome with a known reference. We present a full rewrite of the original pipeline showing a significant reduction of storage requirements by using the CRAM format and runtime by increasing intra-sample parallelization. Both are leading to a 70% cost reduction in commercial clouds enabling users to do large-scale and cross-platform data analysis while keeping costs and CO2 emissions low. The code is available at https://nf-co.re/sarek.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hanssen, F., Garcia, M. U., Folkersen, L., Pedersen, A. S., Lescai, F., Jodoin, S., Miller, E., Wacker, O., Smith, N., nf-core community,, Gabernet, G., Nahnsen, S.. 2023-07-19. Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery. https://doi.org/10.1101/2023.07.19.549462
Cite the original work for its findings. Save a collection to share your selection of sources.