Search bioRxiv⌕ Search

Biology subjects

Taouk, M. L.

Publications and source records attributed to Taouk, M. L..

2 recordsLinked to original sources

Benchmarking long-read simulators against Oxford Nanopore whole-genome sequencing data

BackgroundOxford Nanopore Technologies (ONT) sequencing is increasingly used for whole-genome sequencing (WGS) across a wide range of applications. However, the platform has evolved rapidly through updates to flow cell chemistry and basecalling algorithms, altering the characteristics of the resulting sequencing data. Read simulators provide synthetic datasets with known ground truth, enabling controlled development and evaluation of methods. However, many existing simulators were developed for earlier versions of ONT sequencing or use generic long-read assumptions, and their realism for contemporary ONT data is unclear. ResultsWe benchmarked six ONT-compatible read simulators (Badread, LongISLND, lrsim, NanoSim, PBSIM3 and SimLoRD) using a microbial genome reference and ONT R10.4.1 reads as the empirical standard. Each tool was configured to maximise realism, including training on empirical reads when supported. We compared simulated and real datasets with respect to read length, read accuracy, FASTQ quality scores and sequence error profiles. No simulator reproduced all metrics of the real data well. PBSIM3 most closely reproduced read length, read accuracy and FASTQ quality scores, making it a strong simulator for broad read-level realism. However, it did not capture important features of the real error profile, including context-dependent substitution rates and homopolymer-length errors. Badread and LongISLND better reproduced some aspects of the error profile, but showed other departures from the real data. ConclusionPBSIM3 is a good general-purpose choice for many ONT WGS simulation tasks because it reproduced several key read-level properties well. However, Badread or LongISLND may be preferable for applications where error structure is more important. No evaluated tool was realistic across all tested metrics, highlighting a gap for improved long-read simulators.

bioinformatics↗

Exploring SNP Filtering Strategies: The Influence of Strict vs Soft Core

Phylogenetic analyses are crucial for understanding microbial evolution and infectious disease transmission. Bacterial phylogenies are often inferred from single nucleotide polymorphism (SNP) alignments, with SNPs as the fundamental signal within these data. SNP alignments can be reduced to a strict core by removing those sites which do not have data present in every sample. However, as sample size and genome diversity increase, a strict core can shrink markedly, discarding potentially informative data. Here, we propose and provide evidence to support the use of a soft core that tolerates some missing data, preserving more information for phylogenetic analysis. Using large datasets of Neisseria gonorrhoeae and Salmonella enterica serovar Typhi, we assess different core thresholds. Our results show that strict cores can drastically reduce informative sites compared to soft cores. In a 10,000-genome alignment of Salmonella enterica serovar Typhi, a 95% soft core yielded 10 times more informative sites than a 100% strict core. Similar patterns were observed in Neisseria gonorrhoeae. We further evaluated the accuracy of phylogenies built from strict and soft-core alignments using datasets with strong temporal signals. Soft-core alignments generally outperformed strict cores in producing trees displaying clock-like behaviour; for instance, the Neisseria gonorrhoeae 95% soft core phylogeny had a root-to-tip regression R2 of 0.50 compared to 0.21 for the strict-core phylogeny. This study suggests that soft-core strategies are preferable for large, diverse microbial datasets. To facilitate this, we developed Core-SNP-filter (github.com/rrwick/Core-SNP-filter), an open-source software tool for generating soft-core alignments from whole-genome alignments based on user-defined thresholds. IMPACT STATEMENTThis study addresses a major limitation in modern bacterial genomics - the significant data loss observed in large datasets for phylogenetic analyses, often due to strict-core SNP alignment approaches. As bacterial genome sequence datasets grow and diversity increases, a strict-core approach can greatly reduce the number of informative sites, compromising phylogenetic resolution. Our research highlights the advantages of soft-core alignment methods which tolerate some missing data and retain more genetic information. To streamline the processing of alignments, we developed Core-SNP-filter (github.com/rrwick/Core-SNP-filter), a publicly available resource-efficient tool that filters alignments to informative and core sites. DATA SUMMARYAll genomic sequence reads used in this study were already publicly available and accessions can be found in Supplementary Dataset 1. Supplementary methods and all code can be found in the accompanying GitHub repository: (github.com/mtaouk/Core-SNP-filter-methods).

genomics↗