Search bioRxiv⌕ Search

Biology subjects

Rashid, U.

Publications and source records attributed to Rashid, U..

2 recordsLinked to original sources

nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality

The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.

bioinformatics↗

EDTA v2: enabling scalable TE annotation in animal genomes

The Extensive de-novo TE Annotator (EDTA) automates transposable element annotation in plant genomes but lacks direct LINE/SINE detection, limiting its applicability to animal genomes. We present EDTA v2, which integrates LINE and SINE detection, completely rewrites TIR-Learner for deployability and scalability, and accelerates structural detectors by up to two orders of magnitude. Tested in 30 animal genomes from the Vertebrate Genomes Project Phase I, EDTA v2 bridges the non-LTR detection gap that has prevented automated TE annotation in animals.

genomics↗