Search bioRxiv⌕ Search

Biology subjects

Floden, E. W.

Publications and source records attributed to Floden, E. W..

2 recordsLinked to original sources

An nf-core framework for the systematic comparison of alternative modeling tools: the multiple sequence alignment case study

The computational complexity of many key bioinformatics problems has resulted in numerous alternative heuristic solutions, where no single approach consistently outperforms all others. This creates difficulties for users trying to identify the most suitable tool for their dataset and for developers managing and evaluating alternative methods. As data volumes grow, deploying these methods becomes increasingly difficult, highlighting the need for standardized frameworks for seamless tool deployment and comparison in HPC environments. Multiple sequence aligners (MSAs) rank among the most commonly employed modeling techniques in bioinformatics, playing a crucial role in applications such as protein structure prediction, phylogenetic reconstruction, and variant effect prediction. The NP-hardness of MSAs makes them a major example of problems where heuristics stand central, as no optimal solution can be currently obtained, within the limits of operational computational requirements. Here, we present a pilot design of an nf-core framework for streamlined tool deployment and rigorous performance evaluation focusing on the MSAs software ecosystem. By integrating the most popular MSA tools and focusing on a modular, and extensible architecture, we aspire to provide a key platform supporting MSA deployment, evaluation, and algorithmics development to the MSA community, and a proof-of-principle to the wider bioinformatics community.

bioinformatics↗

Empowering bioinformatics communities with Nextflow and nf-core

Standardised analysis pipelines are an important part of FAIR bioinformatics research. Over the last decade, there has been a notable shift from point-and-click pipeline solutions such as Galaxy towards command-line solutions such as Nextflow and Snakemake. We report on recent developments in the nf-core and Nextflow frameworks that have led to widespread adoption across many scientific communities. We describe how adopting nf-core standards enables faster development, improved interoperability, and collaboration with the >8,000 members of the nf-core community. The recent development of Nextflow Domain-Specific Language 2 (DSL2) allows pipeline components to be shared and combined across projects. The nf-core community has harnessed this with a library of modules and subworkflows that can be integrated into any Nextflow pipeline, enabling research communities to progressively transition to nf-core best practices. We present a case study of nf-core adoption by six European research consortia, grouped under the EuroFAANG umbrella and dedicated to farmed animal genomics. We believe that the process outlined in this report can inspire many large consortia to seek harmonisation of their data analysis procedures.

bioinformatics↗