Search bioRxiv⌕ Search

Biology subjects

Lesack, K.

Publications and source records attributed to Lesack, K..

2 recordsLinked to original sources

The impact of FASTQ and alignment read order on structural variation calling from long-read sequencing data

BackgroundStructural variation (SV) calling from DNA sequencing data has been challenging due to several factors, such as the ambiguity of short-read alignments, multiple complex SVs in the same genomic region, and the lack of "truth" datasets for benchmarking. Additionally, caller choice, parameter settings, and alignment method are known to affect SV calling. However, the impact of FASTQ read order on SV calling has not been explored for long-read data. ResultsIn this study, we used PacBio DNA sequencing data from 15 Caenorhabditis elegans isolates to evaluate the dependence of different SV callers on FASTQ read order. Comparisons of variant call format (VCF) files generated from the original and permutated FASTQ files demonstrated that the order of input data had a large impact on SV prediction, particularly for pbsv. The overall differences were lowest for Sniffles, regardless of the aligner used. The type of variant most affected by read order varied by caller. For pbsv, most differences occurred for deletions and duplications, while for Sniffles, permutating the read order had a stronger impact on insertions. For SVIM, inversions and deletions accounted for most differences. ConclusionThe results of this study highlight the dependence of SV calling on the order of reads encoded in FASTQ files, which has not been recognized in long-read approaches. These findings have implications for the replication of SV studies and the development of consistent SV calling protocols. Our study suggests that researchers should pay attention to the order of reads when analyzing long-read sequencing data for SV calling.

bioinformatics↗

Accurate detection of structural variation is hard

The accurate characterization of structural variation is crucial for our understanding of how large chromosomal alterations affect phenotypic differences and contribute to genome evolution. Whole-genome sequencing is a popular approach for identifying structural variants, but the accuracy of popular tools remains unclear due to the limitations of existing benchmarks. Moreover, the performance of these tools for predicting variants in non-human genomes is less certain, as most tools were developed and benchmarked using data from the human genome. To address this problem, multiple short- and long-read tools were benchmarked using real and simulated Caenorhabditis elegans whole-genome sequence data. To evaluate the use of long-read data for the validation of short-read predictions, the agreement between predictions from a short-read ensemble learning method and long-read tools were compared. The results obtained from simulated data indicate that the best performing tool is contingent on the type and size of the variant, as well as the sequencing depth of coverage. These results also highlight the need for reference datasets generated from real data that can be used as ground truth in benchmarks.

genomics↗