Search bioRxiv⌕ Search

Biology subjects

Pechlivanis, N.

Publications and source records attributed to Pechlivanis, N..

3 recordsLinked to original sources

Synth4bench: a framework for generating synthetic genomics data for the evaluation of tumor-only somatic variant calling algorithms

MotivationSomatic variant calling is a key activity towards identifying genomic alterations; yet, the evaluation of the respective tools remains challenging due to the scarcity of high quality ground truth datasets. To overcome this limitation, we developed synth4bench, a synthetic data generation pipeline for robust benchmarking. Using a systematic process to create distinct synthetic datasets, we thoroughly evaluated five variant callers (Mutect2, FreeBayes, VarDict, VarScan2 and LoFreq). We compared tool outputs against our synthetic ground truth across key sequencing aspects (such as depth and read length) to assess their capacities and shed light on their underlying algorithmic principles. ResultsSynth4bench is an approach for evaluating tumor-only somatic variant callers that relies on a systematic definition of fully controlled ground-truth datasets. Our analysis revealed significant inconsistencies among the tool outputs and a strong dependence of caller performance on sequencing parameters. Indels remain the hardest-to-call variant type, driven by errors at low allele frequencies. Algorithmic choice is also critical; the most robust callers displayed the highest precision in allele frequency estimation, while the most sensitive caller was best for maximizing true positive recovery. Conversely, the least suitable caller exhibited systematic errors along with the poorest overall performance. These findings indicate that there isnt a one-solution-fit-all; sequencing optimization together with caller selection are necessary to maximize sensitivity and reliability. Furthermore, the pronounced inconsistencies suggest that current algorithms are not yet able to capture all mutational mechanisms adequately, with the modeling of the underlying processes remaining an open challenge. Availabilitycode: https://github.com/sfragkoul/synth4bench/ and data: https://zenodo.org/records/16524193 Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=69 SRC="FIGDIR/small/582313v2_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@1ac0fe5org.highwire.dtl.DTLVardef@1479cddorg.highwire.dtl.DTLVardef@8b7d1borg.highwire.dtl.DTLVardef@1c2a81b_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Tracking SARS-CoV-2 variants of concern in wastewater: an assessment of nine computational tools using simulated genomic data

Wastewater-based surveillance (WBS) is an important epidemiological and public health tool for tracking pathogens across the scale of a building, neighbourhood, city, or region. WBS gained widespread adoption globally during the SARS-CoV-2 pandemic for estimating community infection levels by qPCR. Sequencing pathogen genes or genomes from wastewater adds information about pathogen genetic diversity which can be used to identify viral lineages (including variants of concern) that are circulating in a local population. Capturing the genetic diversity by WBS sequencing is not trivial, as wastewater samples often contain a diverse mixture of viral lineages with real mutations and sequencing errors, which must be deconvoluted computationally from short sequencing reads. In this study we assess nine different computational tools that have recently been developed to address this challenge. We simulated 100 wastewater sequence samples consisting of SARS-CoV-2 BA.1, BA.2, and Delta lineages, in various mixtures, as well as a Delta-Omicron recombinant and a synthetic "novel" lineage. Most tools performed well in identifying the true lineages present and estimating their relative abundances, and were generally robust to variation in sequencing depth and read length. While many tools identified lineages present down to 1% frequency, results were more reliable above a 5% threshold. The presence of an unknown synthetic lineage, which represents an unclassified SARS-CoV-2 lineage, increases the error in relative abundance estimates of other lineages, but the magnitude of this effect was small for most tools. The tools also varied in how they labelled novel synthetic lineages and recombinants. While our simulated dataset represents just one of many possible use cases for these methods, we hope it helps users understand potential sources of noise or bias in wastewater sequencing data and to appreciate the commonalities and differences across methods.

bioinformatics↗

Assessing SARS-CoV-2 evolution through the analysis of emerging mutations

IntroThe number of studies on SARS-CoV-2 published on a daily basis is constantly increasing, in an attempt to understand and address the challenges posed by the pandemic in a better way. Most of these studies also include a phylogeny of SARS-CoV-2 as background context, always taking into consideration the latest data in order to construct an updated tree. However, some of these studies have also revealed the difficulties of inferring a reliable phylogeny. [13] have shown that reliable phylogeny is an inherently complex task due to the large number of highly similar sequences, given the relatively low number of mutations evident in each sequence. MotivationFrom this viewpoint, there is indeed a challenge and an opportunity in identifying the evolutionary history of the SARS-CoV-2 virus, in order to assist the phylogenetic analysis process as well as support researchers in keeping track of the virus and the course of its characteristic mutations, and in finding patterns of the emerging mutations themselves and the interactions between them. The research question is formulated as follows: Detecting new patterns of co-occurring mutations beyond the strain-specific / strain-defining ones, in SARS-CoV-2 data, through the application of ML methods. AimGoing beyond the traditional phylogenetic approaches, we will be designing and implementing a clustering method that will effectively create a dendrogram of the involved sequences, based on a feature space defined on the present mutations, rather than the entire sequence. Ultimately, this ML method is tested out in sequences retrieved from public databases and validated using the available metadata as labels. The main goal of the project is to design, implement and evaluate a software that will automatically detect and cluster relevant mutations, that could potentially be used to identify trends in emerging variants. Contacttasos1109@gmail.com

bioinformatics↗