Search bioRxiv⌕ Search

Biology subjects

Colpus, M.

Publications and source records attributed to Colpus, M..

7 recordsLinked to original sources

Long-read sequencing of Mycobacterial tuberculosis is comparable to short-read sequencing for antimicrobial resistance prediction and epidemiological studies.

BackgroundShort-read genetic sequencing technologies (mainly Illumina) have been extensively used for around a decade for Mycobacterium tuberculosis complex (MTBC) outbreak analysis and genomic drug susceptibility testing (gDST) with the result that Illumina has become the de facto gold standard. Long-read sequencing, as exemplified by Oxford Nanopore Technologies (ONT), offer the prospect of faster, simpler, and portable sequencing. In this work, we carry out the largest to date comparison of how well Illumina and ONT technologies sequence MTBC samples, making use of R10.4 ONT flowcells, updated basecalling models and deep-learning variant calling. MethodsA total of 508 samples were sequenced using both short and long-read platforms. All samples originated from South Africa or Vietnam and were over-selected for drug resistance and also included several local outbreaks and a range of lineages. The South African and Vietnamese samples had already been Illumina sequenced. Samples with [≥]50 read depth by Illumina were selected for sequencing by ONT using one of the GridION or PromethION platforms. Bioinformatics processing was done using a modified online cloud platform which included reference-based variant calling, catalogue-based gDST and identified related samples via SNP counting to inform outbreak detection. The lineages and gDST predictions obtained by short-and long-sequencing were compared for all samples as were all putative clusters identified via SNP counting. For convenience Illumina was used as the reference method. FindingsOf the 508 samples, 425 (83.7%) had sufficient read depths to permit comparison between the two sequencing technologies. The assigned lineages were identical for 407/425 (95.8%) samples and all discordances were due to mixed lineages being identified by one technology. Evidence of non-tuberculous mycobacterium (NTM) subpopulations were found in nine samples. Using Illumina as the reference method, the very major error (VME) rate of ONT for predicting resistance to all 15 drugs is 1.0% (0.6-1.5%) whilst the major error (ME) rate is 1.7% (1.3-2.2%) with an unclassified rate of 6.9% (6.3-7.5%). This is below the thresholds specified by the CLSI. Considering each of the 15 drugs individually they had VME and ME point estimates below [≤]3% in 29/30 cases; and most 25/30 below [≤]1.5%. Filtering out all samples containing mixtures left 382 isolates. By appropriate masking of the reference genome we were able to obtain a mean SNP distance between the two platforms of 0.13 (median of zero) for the same sample and for 376/382 samples (98.4%, CI:96.6-99.4%) the difference was [≤]1 SNPs. The high concordance in SNP identification ensured that few differences in the 43 putative clusters among 172 isolates were observed. InterpretationThe differences between the two sequencing platforms for the key clinical outputs is so small that it is now within the tolerances set by regulatory agencies. Provided the sequencing is of sufficient quality, we have therefore reached a threshold whereby sequencing data from long-and short-read platforms can be aggregated. This will enable large scale analyses by national and international public health agencies whilst allowing the MTBC community to take advantage of the portability and speed of long-read sequencing. FundingThe NIHR Health Protection Research Unit: Healthcare Associated Infections and Antimicrobial Resistance at University of Oxford (NIHR200915), a partnership between the UK Health Security Agency (UKHSA) and the University of Oxford, the National Institute for Health and Care Research Biomedical Research Centre: Oxford (BRC) and the Ellison Institute of Technology, Oxford Ltd. The CRyPTIC project was funded by Wellcome [214560/Z/18/Z], a Wellcome Trust/Newton Fund-MRC Collaborative Award (200205/Z/15/Z); and the Bill & Melinda Gates Foundation Trust (OPP1133541). Research in contextO_ST_ABSEvidence before this studyC_ST_ABSWe conducted a PubMed Central full text search for "tuberculosis" AND ("drug resistance prediction" OR "drug susceptibility prediction") AND ("genome" OR "genomic" OR "geno-typic") AND ("ont" OR "oxford nanopore") between 2022 and 2026 (conducted 1 April 2026). This returned 62 papers; of which, six used both Illumina and ONT sequencing. One of these, published in 2023, directly compared the performance of the two platforms on 151 M. tuberculosis isolates oversampled for resistance. The investigation yielded comparative results for the earlier generation ONT flow cell (R9{middle dot}4{middle dot}1) and base-caller (guppy version 5{middle dot}0{middle dot}16). Another, published in 2026, investigated a targeted next-generation sequencing panel of 20 amplicons using ONT sequencing on R10.4.1 flow cells with guppy 6{middle dot}4{middle dot}6. They compared the results on 71 isolates against phenotypic data and Illumina whole genome sequencing (for 53 isolates) but had low rates of resistance, with all drugs but isoniazid being limited to under five resistant isolates. Two other small studies (10 and 13 samples, respectively) conducted feasibility studies comparing ONT with Illumina, also using earlier generation flow cells and base-calling technology from ONT. Two further studies compared Illumina with ONT for direct sputum sequencing and did not investigate the comparative performance of the two platforms for variant call accuracy, resistance prediction, and outbreak detection. Illumina sequencing technology is widely used for genomic sequence analysis in research, and clinical and public health contexts. Consequently, it has become the de facto reference standard for generating whole genome sequence data. Whilst previous studies established the promise and limitations of long-read (ONT) sequencing as an alternative to short-read sequencing (mainly Illumina), the enhanced performance arising from newer flowcells (e.g. R10.4.1), V14 chemistry, and the latest basecallers (dorado v4.3.0/5.0.0) has not been analysed. Neither has any ONT analysis incorporating the new deep-learning variant callers been evaluated in a large-scale comparative study. Thus, it is currently unclear whether data generated by either platform can be used safely in aggregated analyses for research and clinical or public health service. Added value of this studyWe compared how well short-(Illumina) and long-read (ONT) sequencing platforms identify the genetic variants in M. tuberculosis, predict antituberculous drug resistance and recog-nise outbreaks. The long-reads were generated using the latest generation ONT R10.4.1 flows cells, V14 chemistry, super high accuracy basecalling (dorado v4.3.0/5.0.0) and a bioinformatics analysis pipeline built using the Clair3 deep-learning based variant caller. A total of 508 clinical samples were sequenced using both technologies, substantially more than previous studies. The sampling frame was much larger than previously investigations and included a large proportion of isolates with resistance to first-line and second-line antibiotics as well as bedaquiline. Thus, providing greater statistical power for resistance prediction than before. In particular, the inclusion of bedaquiline resistance provided evidence useful for predicting resistance to this newly deployed drug for treating multi-drug resistant (MDR) TB. We find that the differences between technologies are small meaning that either technology can be used alone safely, and services using both technologies can confidently aggregate the data for analysis. Implications of all the available evidenceThis will be a benefit to local, regional and international organisations, particularly public health agencies, which often have a mix of the two main sequencing technologies for characterising TB whole genome sequences. It also opens up the sequence based diagnostic market to greater competition, particularly if the observed performance can be replicated for other pathogen species.

microbiology↗

OxBreaker: species-agnostic pipeline for the analysis of outbreaks using nanopore sequencing

Real-time genomic surveillance may mitigate the spread of health-care-associated infections, but whole-genome sequencing costs and the need for specialised expertise constrain its wide implementation in public health. Here we present OxBreaker, an automated and species-agnostic pipeline optimised for the high-resolution analysis of bacterial and plasmid genomes sequenced via Oxford Nanopore Technologies (ONT). OxBreaker streamlines the transition from raw reads to phylogenetic inference through automated reference selection and high-accuracy variant calling. It is accessible via a graphical user interface (GUI) that can be easily installed locally and operated by non-specialists. Benchmarking against technical and biological replicates of high-priority pathogens demonstrates high accuracy, with false positive variant rates reduced to 0-4 single-nucleotide polymorphisms (SNPs) for common species. We further validated the pipeline by accurately characterising previously published clonal and plasmid-mediated outbreaks, reproducing established phylogenies with improved accessibility. By providing a stable, scalable, open-source offline-compatible solution that matches the resolution of short-read platforms while maintaining the speed of long-read technology, OxBreaker is designed to facilitate the adoption of local, real-time genomic surveillance for frontline infection prevention and control.

genomics↗

Evaluation of an Oxford Nanopore sequencing workflow for mycobacteria from primary MGIT culture

Illumina sequencing of primary MGIT cultures is an established workflow in several reference mycobacteriology laboratories. Oxford Nanopore Technologies (ONT) provides real-time genetic sequencing yielding long reads which help resolve repetitive genomes and is being explored for in-house implementation within diagnostic laboratories. However, low DNA yields from primary MGIT cultures frequently limit the application of ONT workflows, due to high minimum DNA input requirements for library preparation. We validated a modified ONT workflow combining rapid, semi-automated DNA extraction from MGIT cultures with Rapid PCR-Barcoding for whole-genome amplification, and compared its performance with Illumina sequencing for species identification and Mycobacterium tuberculosis complex (MTBC) single-nucleotide polymorphism (SNP) detection. A platform-agnostic analysis pipeline enabled consistent human read removal, taxonomic assignment, and MTBC genomic characterisation. ONT sequencing data was subsampled at 1, 6, and 72 hours to determine the earliest time point for reliable species identification. The concordance between sequencing platforms on species classification was 98.3% (95.8-99.5%) with all differences arising from potential mixed infections. SNP agreement was high, with a mean of 0.3 and a median of 0 SNP differences between sequencing platforms after masking. These findings demonstrate the feasibility of PCR-amplified ONT sequencing as a reliable alternative for routine genomic characterisation of MGIT cultures. IMPORTANCERapid identification of mycobacterial infections is essential for timely patient care and infection control. Many clinical laboratories currently rely on outsourcing sequencing to external reference centres, which adds time and delays the return of clinically actionable results. These services commonly use Illumina sequencing, which, while highly accurate, involves complex workflows and longer turnaround times. Newer technologies offer the potential to generate results in real time, but their use has been limited by the low amount of DNA available from routine culture samples. In this study, we developed an improved workflow that increases the amount of usable DNA and enables reliable, rapid sequencing directly from these samples. This approach allows multiple samples to be processed together and reduces the time needed to obtain results. Importantly, it could enable clinical laboratories to perform sequencing in-house, reducing reliance on external services and improving turnaround times.

molecular biology↗

Rapidly and reproducibly building a comprehensive catalogueof resistance-associated variants for M. tuberculosis

BackgroundCatalogues of genetic variants associated with resistance underpin whole-genome sequencing (WGS)-based predictions of drug susceptibility in Mycobacterium tuberculosis, and are essential for molecular diagnostics and surveillance. The current gold standard catalogues are those released by the WHO but the underlying data are not fully released and they are difficult to interpret. Open and reproducible methods would help address these problems, extending the important work already done. MethodsWe have developed an automated method, catomatic, that uses a binomial test to associate informative isolates with resistance or susceptibility, and built a catalogue (catomatic-1) from the same 39,358 samples used to construct the first edition of the WHO catalogue (WHOv1). We performed a sensitivity analysis to optimise statistical and bioinformatic parameters for each drug, and benchmarked catomatic-1 against WHOv1 using an independent Validation Dataset of 14,380 isolates. FindingsBy using simpler statistics, catomatic-1 algorithmically classified 1,329 genetic variants, ranging from five for linezolid to 440 for pyrazinamide. WHOv1 included generalisable rules added by a panel of experts, increasing its predictive coverage, but at the cost of reproducibility. Despite not including such expert rules, catomatic-1 achieves comparable performance for all drugs, with sensitivities for first-line agents above 88% on the independent Validation Dataset. The automated process allowed us to efficiently explore parameter space; for instance, detecting resistant variants with low read support improved the sensitivity for all drugs. InterpretationPerformant resistance catalogues for M. tuberculosis can be built automatically using transparent and reproducible statistical methods. As more data are collected, catalogue content and performance will evolve, highlighting the need for proper versioning, machine/human readability, and open access. This approach demonstrates resistance catalogues used in surveillance and diagnostics can be rapidly and reproducibily updated. FundingThe National Institute for Health and Care Research (NIHR), Engineering and Physics Sciences Research Council (EPSRC) and ORACLE Corporation. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed and preprint servers (bioRxiv, medRxiv), and publicly available mutation catalogues for studies linking Mycobacterium tuberculosis genomic variants with drug resistance using whole-genome or targeted sequencing and phenotypic drug-susceptibility testing (pDST). Search terms combined "Mycobacterium tuberculosis", "genome sequencing", "mutation catalogue", "mutation effects", "drug resistance", and individual drug names, with no language or date restriction. We included studies providing paired, clinical genomic and pDST or MIC data, excluding purely in-silico or case-only reports. This work directly builds on methodologies and data published by five prior studies, and makes primary comparisons with the First (WHOv1) and Second (WHOv2) Editions of the WHO Catalogue of mutations in Mycobacterium tuberculosis. Added value of this studyWe developed catomatic, a transparent, reproducible tool for building catalogues of resistance- and susceptibility-associated genetic variants. Trained on the same samples used to build WHOv1 and benchmarked on an independent Validation Dataset, catomatic achieves comparable sensitivity, specificity, and definitive prediction rates to WHOv1 without expert-rule augmentation and despite using simpler statistics. It optimises parameters per drug, produces machine-readable outputs (CSV/JSON), and demonstrates that adjusting read-support thresholds can improve detection of minor resistance subpopulations. Implications of all the available evidenceCatalogues of resistance-associated variants for M. tuberculosis can be rapidly and transparently constructed. Making catalogues available in human/machine-readable formats with uncertainty estimates will improve uptake of WGS for M. tuberculosis surveillance and diagnostics; using a reproducible process permits diagnostic test manufacturers, researchers, clinical and public health laboratories to select the level of statistical support necessitated by their specific use-case, Policymakers should balance the benefits of expert rules against loss of reproducibility. Future work will expand the size of the datasets used, integrate minimum inhibitory concentration data, and establish consensus workflows for routine, transparent catalogue updates.

microbiology↗

Enhancement and validation of the antibiotic resistance prediction performance of a cloud-based genetics processing platform for Mycobacteria

Tuberculosis remains a global health problem. Making it easier and quicker to identify which antibiotics an infection is likely to be susceptible to will be a key part of the solution. Whilst whole-genome sequencing offers many advantages, the processing of the genetic reads to produce the relevant public health and clinical information is, surprisingly, often the responsibility of the end user which inhibits uptake. Here we characterise how well a freely-available tool we have developed, gnomonicus, predicts the antibiotic resistance profile of a sample (given its variant call file) using our implementation of the second edition of the WHO catalogue of resistance-associated variants (WHOv2). To facilitate this, we have constructed a Diverse Testset of 2,663 publicly-available M. tuberculosis samples which have both genetic and drug susceptibility testing (DST) data. We have chosen to apply the catalogue such that our tool will return a result of (i) Fail if there are insufficient reads at a genetic locus associated with resistance, (ii) Unknown if a genetic variant in a resistance gene not listed in the catalogue is encountered and (iii) Resistant if three or more short-reads support the presence of a resistance-associated variant. The last step increases the sensitivity for all 15 antibiotics but only reaches significance in a few in our testset. Comparing our results to those of TB-Profiler, an existing tool, highlights the different design choices and demonstrates the performance of both tools on our Diverse Testset is comparable. By only considering high confidence DST results we show that gnomonicus, in combination with our translation of WHOv2, achieves sensitivities and specificities in excess of 95% for both isoniazid and rifampicin. Impact StatementWhole genome sequencing clinical samples taken from patients with tuberculosis is a potentially fast and accurate method for determining to which antibiotics the infection will be susceptible. Two barriers need to be overcome; the first, which is knowing which mutations are associated with resistance (or not) to a range of antibiotics is well on the way to be solved thanks to the efforts of the World Health Organization (WHO) who have published extensive catalogues containing lists of such mutations. The second barrier is that the processing of the raw genetic files remains a largely manual process overseen by bioinformaticians. Here we describe gnomonicus, our open-source AMR prediction tool, and report the performance of our translation of the second edition of the WHO catalogue using a carefully designed publicly-available dataset of 2,663 M. tuberculosis samples. We hope that not only will this tool be useful but also that this dataset will be used by other researchers to facilitate comparisons between pipelines, approaches and tools. Data SummaryThe attendant GitHub repository1 allows gnomonicus to be rerun on all 2,663 samples in the Diverse Testset; it therefore includes instructions, the necessary input files (all the variant call files, the version of the WHOv2 catalogue and a link to the H37Rv GenBank file used in this study) and also the output JSON files. The ENA accession numbers for all 2,663 samples (including a bash script to download them) and their corresponding phenotypic drug susceptibility testing results are included with the intention that people can either reproduce our results, or use the same dataset for other analyses. The JSON files output by TB-Profiler have also been added for comparison. The repository contains a series of Juypter notebooks containing Python3 code that allows the user to discover and parse the output JSON files from either tool and save the results as data tables. Other notebooks allow the user to reproduce all the analysis underlying this work, including reproducing the figures and many of the tables.

microbiology↗

E. coli phylogeny drives co-amoxiclav resistance through variable expression of blaTEM-1

Co-amoxiclav resistance in E. coli is a clinically important phenotype associated with increased mortality. The class A beta-lactamase blaTEM-1 is often carried by co- amoxiclav-resistant pathogens, but exhibits high phenotypic heterogeneity, making genotype-phenotype predictions challenging. We present a curated dataset of n=377 E. coli isolates representing all 8 known phylogroups, where the only acquired beta- lactamase is blaTEM-1. For all isolates, we generate hybrid assemblies and co-amoxiclav MICs, and for a subset (n=67/377), blaTEM-1 qPCR expression data. First, we test whether certain E. coli lineages are intrinsically better or worse at expressing blaTEM-1, for example, due to lineage differences in regulatory systems, which are challenging to directly quantify. Using genotypic features of the isolates (blaTEM-1 promoter variants and copy number), we develop a hierarchical Bayesian model for blaTEM-1 expression that controls for phylogeny. We establish that blaTEM-1 expression intrinsically varies across the phylogeny, with some lineages (e.g. phylogroups B1 and C, ST12) better at expression than others (e.g. phylogroups E and F, ST372). Next, we test whether phylogenetic variation in expression influences the resistance of the isolates. With a second model, we use genotypic features (blaTEM-1 promoter variants, copy number, duplications; ampC promoter variants; efflux pump AcrF presence) to predict isolate MIC, again controlling for phylogeny. Lastly, we use a third model to demonstrate that the phylogenetic influence on blaTEM-1 expression causally drives the variation in co- amoxiclav MIC. This underscores the importance of incorporating phylogeny into genotype-phenotype predictions, and the study of resistance more generally.

microbiology↗

Evaluation of the accuracy of bacterial genome reconstruction with Oxford Nanopore R10.4.1 long-read-only sequencing

2.Whole genome reconstruction of bacterial pathogens has become an import tool for tracking antimicrobial resistance spread, however accurate and complete assemblies have only been achievable using hybrid long and short-read sequencing. We have previously found the Oxford Nanopore Technologies (ONT) R10.4/kit12 flowcells produced improved assemblies over the R9.4.1/kit10, however they contained too many errors compared to hybrid Illumina-ONT assemblies. ONT have since released the R10.4.1/kit12 flowcells that promises greater accuracy and yield. They have also released newly trained basecallers using native bacterial DNA containing methylation sites intended to fix systematic errors, specifically Adenosine (A) to Guanine (G) and Cytosine (C) to Thymine (T) substitutions. ONT have recommended the use of Bovine Serum Albumin (BSA) during library preparation to improve sequencing yield and accuracy. To evaluate these improvements, we sequenced DNA extracts from four commonly studied bacterial pathogens, namely Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa and Staphylococcus aureus, as well as 12 disparate E. coli clinical samples from different phylogroups and sequence types. These were all sequenced with and without BSA. These sequences were de novo assembled and compared against Illumina-corrected reference genomes. Here we have found the nanopore long read-only R10.4.1 (kit14) assemblies with basecallers trained using native bacterial methylated DNA produce accurate assemblies from 40x depth or higher, sufficient to be cost-effective compared to hybrid long-read (ONT) and short-read (Illumina) sequencing. 3. Impact statementCurrently, the best method of building accurate and complete bacterial genome assemblies is to create a hybrid assembly; combining both long and short DNA sequencing reads. Short reads are more accurate, but can be difficult to assemble into a complete genome. Long reads are generally less accurate, but easier to reconstruct into a complete genome. By combining long and short reads, we get both accuracy and reconstructive power. However, this also involves higher costs and more labour than using a single sequencing platform. In this study, we compare long read only assemblies from Oxford Nanopore Technologys newest iteration of improvements in both chemistry and software to hybrid Illumina-Nanopore assemblies. We sequenced four bacterial pathogens with published reference genomes (Staphylococcus aureus, Klebsiella Pneumoniae, Pseudomonas Aeruginosa, and Escherichia Coli) and twelve bloodstream associated E. coli, and show that assemblies from the newest technology are not only an improvement on the previous iteration, but are able to compete with hybrid Illumina-Nanopore assemblies in their quality, providing a step towards bacterial genome assembly using a single sequencing platform. 4. Data summaryThe authors confirm all supporting data, code and protocols have been provided within the article, through supplementary data files, or in publicly accessible repositories. Nanopore and Illumina fastq data are available in the ENA under project accession: PRJEB51164. Assemblies have been made available at: https://figshare.com/articles/dataset/R10_4_1_KIT14_comparison_assemblies/2497 2954 Code and analysis outputs are available at: https://gitlab.com/ModernisingMedicalMicrobiology/assembly_comparison

microbiology↗