Search bioRxiv⌕ Search

Biology subjects

Featherstone, L. A.

Publications and source records attributed to Featherstone, L. A..

3 recordsLinked to original sources

Exploring SNP Filtering Strategies: The Influence of Strict vs Soft Core

Phylogenetic analyses are crucial for understanding microbial evolution and infectious disease transmission. Bacterial phylogenies are often inferred from single nucleotide polymorphism (SNP) alignments, with SNPs as the fundamental signal within these data. SNP alignments can be reduced to a strict core by removing those sites which do not have data present in every sample. However, as sample size and genome diversity increase, a strict core can shrink markedly, discarding potentially informative data. Here, we propose and provide evidence to support the use of a soft core that tolerates some missing data, preserving more information for phylogenetic analysis. Using large datasets of Neisseria gonorrhoeae and Salmonella enterica serovar Typhi, we assess different core thresholds. Our results show that strict cores can drastically reduce informative sites compared to soft cores. In a 10,000-genome alignment of Salmonella enterica serovar Typhi, a 95% soft core yielded 10 times more informative sites than a 100% strict core. Similar patterns were observed in Neisseria gonorrhoeae. We further evaluated the accuracy of phylogenies built from strict and soft-core alignments using datasets with strong temporal signals. Soft-core alignments generally outperformed strict cores in producing trees displaying clock-like behaviour; for instance, the Neisseria gonorrhoeae 95% soft core phylogeny had a root-to-tip regression R2 of 0.50 compared to 0.21 for the strict-core phylogeny. This study suggests that soft-core strategies are preferable for large, diverse microbial datasets. To facilitate this, we developed Core-SNP-filter (github.com/rrwick/Core-SNP-filter), an open-source software tool for generating soft-core alignments from whole-genome alignments based on user-defined thresholds. IMPACT STATEMENTThis study addresses a major limitation in modern bacterial genomics - the significant data loss observed in large datasets for phylogenetic analyses, often due to strict-core SNP alignment approaches. As bacterial genome sequence datasets grow and diversity increases, a strict-core approach can greatly reduce the number of informative sites, compromising phylogenetic resolution. Our research highlights the advantages of soft-core alignment methods which tolerate some missing data and retain more genetic information. To streamline the processing of alignments, we developed Core-SNP-filter (github.com/rrwick/Core-SNP-filter), a publicly available resource-efficient tool that filters alignments to informative and core sites. DATA SUMMARYAll genomic sequence reads used in this study were already publicly available and accessions can be found in Supplementary Dataset 1. Supplementary methods and all code can be found in the accompanying GitHub repository: (github.com/mtaouk/Core-SNP-filter-methods).

genomics↗

Clockor2: Inferring global and local strict molecular clocks using root-to-tip regression

Molecular sequence data from rapidly evolving organisms are often sampled at different points in time. Sampling times can then be used for molecular clock calibration. The root-to-tip (RTT) regression is an essential tool to assess the degree to which the data behave in a clock-like fashion. Here, we introduce Clockor2, a client-side web application for conducting RTT regression. Clockor2 uniquely allows users to quickly fit local and global molecular clocks, thus handling the increasing complexity of genomic datasets that sample beyond the assumption homogeneous host populations. Clockor2 is efficient, handling trees of up to the order of 104 tips, with significant speed increases compared to other RTT regression applications. Although clockor2 is written as a web application, all data processing happens on the client-side, meaning that data never leaves the users computer. Clockor2 is freely available at https://clockor2.github.io/.

evolutionary biology↗

Assessing the effects of date and sequence data in phylodynamics

Despite its increasing role in the understanding of infectious disease transmission at the applied and theoretical levels, phylodynamics lacks a well-defined notion of ideal data and optimal sampling. We introduce a formal method to visualise and quantify the relative impact of pathogen genome sequence and sampling times--two fundamental sources of data for phylodynamics under birth-death-sampling models--to understand how each drive phylodynamic inference. Applying our method to simulations and outbreaks of SARS-CoV-2 and H1N1 Influenza data, we use this insight to elucidate fundamental trade-offs and guidelines for phylodynamic analyses to draw the most from sequence data. Phylodynamics promises to be a staple of future responses to infectious disease threats globally. Continuing research into the inherent requirements and trade-offs of phylodynamic data and inference will help ensure phylodynamic tools are wielded in ever more targeted and efficient ways.

evolutionary biology↗