Search bioRxiv⌕ Search

Biology subjects

Anker, K. M.

Publications and source records attributed to Anker, K. M..

2 recordsLinked to original sources

Identification and Masking of Artefactual and Misleading Within-Host Variants in Deep-Sequencing SARS-CoV-2 Data

Deep sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artefacts. In this study, we show that recurrent artefactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency (MAF) thresholds. Using data from the UKs Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artefacts are predominantly sequencing centre-rather than protocol-specific. Each centre exhibits a modest, distinct set of recurrent artefactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artefactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Together, these findings highlight the importance of explicit, dataset-aware artefact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens such as SARS-CoV-2.

genomics↗

Exploring genetic signatures of zoonotic influenza A virus at the swine-human interface with phylogenetic and ancestral sequence reconstruction

Influenza A viruses (IAVs) in swine have zoonotic potential and pose a continuous threat of causing human pandemics, as demonstrated by the H1N1 pandemic in 2009. Despite increased genomic surveillance, we have limited knowledge of the IAV evolutionary dynamics leading to such zoonotic events and no clear understanding of genetic markers associated with interspecies transmission of IAV between humans and swine. To explore this, we analyzed a comprehensive publicly available whole genome dataset of human and swine IAV sequences. We conducted phylogenetic analyses and inference of ancestral host and sequence states for each IAV segment and mapped inferred mutations that were associated with transmission within and between swine and human hosts. We developed a custom python library to combine information from host and ancestral sequence annotated trees and applied statistical models to identify genetic markers associated with intra- or interspecies transmissions between swine and humans. This included analyzing mutation rates and the selective pressures acting on the viral proteins following intra- and interspecies transmissions and using a scalable, gradient-boosted decision tree machine learning approach to predict key amino acid positions critical for different transmission types. Our analyses indicated complex mutational patterns within and across viral proteins, but also suggested that specific protein regions and amino acid positions of especially several of the internal gene segments were more important for interspecies transmission. Our findings identify potential genetic signatures across the IAV proteins associated with host adaptation and zoonotic potential, offering valuable markers for early-warning genomic surveillance systems to enhance animal health and minimize the potential for zoonotic transmission of IAV.

bioinformatics↗