Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.09.15.676237

Nanopore long-read only genome assembly of clinical Enterobacterales isolates is complete and accurate

Abstract

Whole bacterial genome sequence reconstruction using Oxford Nanopore Technologies ("Nanopore") long-read only sequencing may offer a lower-cost, higher-throughput alternative for pathogen surveillance to hybrid assembly with recent improvements in Nanopore sequencing accuracy. We evaluated the accuracy, including plasmid reconstruction, of Nanopore long-read only genome assemblies of Enterobacterales. We sequenced 92 genomes from clinical Enterobacterales isolates, collected in England under a national surveillance program, with long-read Nanopore (R10.4.1, Dorado v5.0.0 super-high-accuracy basecalled) and short-read Illumina (NovaSeq) sequencing approaches. Genomes were assembled using three long-read only (Flye; Hybracter long; Autocycler), and three hybrid assemblers (Hybracter hybrid; Unicycler normal; bold). Three polishing modalities (Medaka v2 with subsampled or un-subsampled long-reads; Polypolish + Pypolca with short-reads) were investigated. Autocycler circularised the most chromosomes (87/92 [95%]). Plasmid sequence reconstruction was comparable between all assemblers except Flye, all recovering 90-96% of plasmids, although the ground truth was uncertain. Flye performed worse than other assemblers on almost all metrics. Autocycler + Medaka (un-subsampled long-reads) was the most accurate long-read only assembler/polisher combination, comparable to hybrid assemblies (median 0 [IQR:0-0] SNPs and 0 [IQR:0-1] indels per genome; quality value/Q score, 100 [IQR: 64-100]), with only 4/92 genome sequences having >10 SNPs/indels. Medaka polishing with un-subsampled long-reads resulted in small improvements in indels but not SNPs for both Flye and Autocycler assemblies. Seven-locus MLST, antimicrobial resistance, virulence, and stress gene annotation was equivalent across assembler/polisher combinations. Nanopore long-read only bacterial genome assembly with Autocycler combined with Medaka polishing (using un-subsampled reads) is similarly accurate and possibly more complete than hybrid assemblies, representing a viable alternative for incorporating high-quality genomic data, including plasmids, into Enterobacterales surveillance. Data SummaryNanopore long-reads and Illumina short-reads from the 92 Enterobacterales isolates from this study have been uploaded to ENA (BioProject accession: PRJEB93885). Code for the Nextflow assembly pipeline, downstream analysis scripts, and R statistical analysis scripts are available on GitHub (https://github.com/oxfordmmm/NEKSUS_ont_hybrid_assembly_comparison). The following supplementary data tables are available on FigShare (https://figshare.com/account/home#/projects/253775): O_LIENA Sample accessions and sample metadata (accessions_and_metadata.csv) C_LIO_LISeqkit stats summaries of the Illumina and Nanopore reads (raw_qc_sup.cav) C_LIO_LISummary of assembly contig features (contigs_summary_sup_cleaned.csv) C_LIO_LIPairwise mash distances between contigs (mash_cleaned.csv) C_LIO_LIPlasmids matching across different assemblers compared to the Hybracter (hybrid) and manually-curated reference sets (plasmids_match_hybracter_mash.csv; plasmids_match_manual_mash.csv, respectively) C_LIO_LISeven-locus multi-locus sequence type annotation (mlst_cleaned.csv) C_LIO_LICheckM2 summaries of assemblies (checkm2_cleaned.csv) C_LIO_LINucleotide-level accuracy of assemblies (SNP, Indels, and Quality value compared to short-read mapping; assembly_nucleotide_accuracy_cleaned.csv) C_LIO_LIBakta annotation (bakta_by_contig_cleaned.csv) C_LIO_LIAMRFinderPlus annotations of contigs (amrfinder_plus_cleaned.csv) C_LIO_LIMOB-suite annotation summaries of contigs (mobsuite_cleaned.csv) C_LI Impact StatementNanopore long-reads have historically been too error-prone to use alone for accurate bacterial genome assembly, necessitating additional Illumina short-reads to achieve structurally complete and accurate hybrid genome assemblies for public health surveillance. This increases cost and complexity. Previous studies have shown that recent improvements in Nanopore chemistry (R10.4.1 flowcell) and basecalling (super-high accuracy) allow high-quality long-read only assemblies on a small number of laboratory reference strains. This is the first evaluation, to our knowledge, to assess Nanopore long-read only genome assembly compared with hybrid assembly on a large number of clinical isolates. In addition, this is the first large-scale evaluation of the recently released automated consensus long-read assembly tool, Autocycler. We show that Autocycler long-read only assemblies are more structurally complete for chromosomal sequences, while reconstructing a similar number of plasmids to other long-read and hybrid assemblers. Most long-read polished, Autocycler-assembled genome sequences have 0 errors (median: 0 SNPs/indels) relative to a short-read polished (hybrid) Autocycler assemblies, enabling accurate annotation of key genes.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nagy, D., Pennetta, V., Rodger, G., Hopkins, K., Jones, C. R., The NEKSUS Consortium,, Hopkins, S., Crook, D., Walker, A. S., Robotham, J., Hopkins, K. L., Ledda, A., Williams, D., Hope, R., Brown, C. S., Stoesser, N., Lipworth, S.. 2025-09-17. Nanopore long-read only genome assembly of clinical Enterobacterales isolates is complete and accurate. https://doi.org/10.1101/2025.09.15.676237

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

AF3 Inspector: In-Browser 3D Model Visualization and Confidence Analytics for AlphaFold 3

AlphaFold 3 (AF3) has broadened the computational structural biology landscape by predicting all-atom complexes across proteins, nucleic acids, small-molecule ligands, and ions, including chemical modifications. For non-experts, Google provides a fully web-based service to utilize AlphaFold 3. Unfortunately, despite a flexible and rich interface for inputs, the server's outputs are incomplete and complicated for many users, as witnessed hands-on. Not all confidence metrics are displayed, and those that are shown do not present much interactivity; besides, only model number 1 is shown among the 5 produced by the software, and upon download the user finds the models are provided only in CIF format, which is not yet well-known especially among highly practical users. Here, we present the AlphaFold 3 Multi-Model Inspector (AF3 Inspector, https://pdbms.altervista.org/af3viewer/afviewer8.html), an open, client-side, zero-install web application designed to dissect, align, and interactively explore the complete ensemble of structural models produced by AlphaFold 3. Part of the PDB Manipulation Suite (https://pdbms.altervista.org/), the AF3 Inspector operates entirely within web browsers, meaning it is available out of the box in all devices and operating systems. AF3 Inspector delivers 3D visualization in various styles and colors for all models, overlay with backbone alignment if requested, interactive plots for pLDDT profiles, pAE matrices and contact maps, chain-pair interaction heatmaps, automated detection of ligands, ions and post-translational modifications, and rapid mmCIF-to-PDB conversion allowing users to download the more familiar files. The platform addresses critical practical considerations highlighted in recent community benchmarks, such as CASP16, and extends the lightweight, client-side biophysical toolkit paradigm established by the PDB Manipulation Suite (PDBMS).

bioinformatics↗

PanVasc Research for AI assisted evidence analysis in panvascular intervention

Panvascular intervention research requires evidence workflows that preserve source identity, outcome definitions and observation windows. We developed PanVasc Research, an executable research framework, and evaluated a fixed local Qwen3-4B model using two complementary tasks. Fifty ClinicalTrials.gov records from five vascular query strata generated 200 source-fidelity tests with intact evidence or controlled removal of the requested source, primary outcome or timeframe, plus 50 clean controls. A separate 100-statement sample from the official NLI4CT test set assessed clinical-trial entailment and evidence selection against existing expert labels. Generic and checklist prompts used identical evidence and maximum generation budgets; each output was also evaluated with an input-only deterministic contract. Registry exact accuracy was 51/200 (25.5%) with the generic prompt and 77/200 (38.5%) with the checklist; intact-case agreement was 50/50 and 48/50. The rule baseline recovered 200/200 tuples. NLI label accuracy was 48/100 (48.0%) and 50/100 (50.0%), respectively. The source contract retained 35 and 31 incorrect NLI labels in the two arms. Registry references establish fidelity to a registration snapshot, while NLI4CT concerns breast-cancer trials and does not validate vascular expertise. The framework also retains provenance-recorded literature retrieval, structured research planning and local numerical analysis. These experiments support a bounded assessment of source handling and semantic failure, rather than a new foundation model or autonomous scientific discovery. A proposed endpoint representation identifies the additional domain annotation and independent validation required for a panvascular research model.

bioinformatics↗

BiomiX 3.0: A user-friendly platform for democratized multi-omics integration with graph-based learning.

Background Multi-omics integration has emerged as a powerful strategy to decode the molecular complexity of biological systems. However, the diversity of available methods, each designed with distinct assumptions, objectives, and computational requirements, makes method selection, usage and interpretation challenging for nonexpert users. Here we present BiomiX 3.0, an updated version of the BiomiX platform that extends its integration capabilities with four additional methods: Similarity Network Fusion (SNF), NEighborhood-based Multi-Omics clustering (NEMO), Data Integration Analysis for Biomarker discovery using Latent variable approaches for Omics studies (DIABLO), and PRAMIGO (Phenotyping netwoRk Application for Multi-omics InteGratiOn), a novel supervised heterogeneous graph transformer (HGT) introduced in this work. Results We benchmarked all five methods, MOFA, DIABLO, SNF, NEMO, and PRAMIGO, on two independent multiomics datasets derived from a Chronic Lymphocytic Leukemia (CLL) cohort comparing IGHV-mutated and unmutated patients, and a pulmonary tuberculosis (PTB) cohort versus healthy controls. Supervised methods (DIABLO, PRAMIGO) consistently achieved higher condition-specific discrimination as measured by the Adjusted Rand Index (ARI) and the Adjusted Mutual Information (AMI). In contrast, unsupervised methods (SNF, NEMO) revealed alternative patient stratifications driven by independent sources of biological variance while MOFA performed in a semi-supervised way occupies an intermediate position, capturing latent factors that explain both disease-associated and orthogonal sources of variance. Systematic gene-centric analysis of the top-ranked features prioritized by each method was supported by manual biological annotation of shared and method-specific signals. Across cohorts, we annotated 154 shared features (116 genes, 38 metabolites) and 60 method-unique features per cohort, demonstrating that no single integration strategy captures the full landscape of biologically relevant signals. In the CLL cohort, shared features spanned B-cell receptor biology, innate immune signaling, RAS/MAPK activation, and epigenetic regulation, while methodunique features revealed supervised-method-specific insights into vesicle trafficking (DIABLO), immune checkpoints (MOFA), and ncRNA regulation (PRAMIGO). In the PTB cohort, a convergent interferon/innate immune signature dominated shared features across all methods. Still, method-unique analysis uncovered DIABLO-specific acylcarnitine metabolic reprogramming, MOFA-specific restoration of lysosomal trafficking, and PRAMIGO-specific {gamma}{delta} T-cell and immunoglobulin repertoire diversity. Across both cohorts, SNF and NEMO proved useful for detecting biological and technical sources of variation that were orthogonal to the primary condition of interest. Ultimately, PRAMIGO uniquely enables the construction of heterogeneous graphs modeling cross-modal molecular interactions, uncovering epigenetic co-regulation programs in CLL and multi-omics inflammatory modules in PTB that are difficult to identify using conventional integration approaches. Conclusions BiomiX 3.0 provides a graphical user interface (GUI) multi-method integration environment that democratizes access to state-of-the-art multi-omics analysis. By combining both unsupervised and supervised integration strategies within a unified platform and introducing graph-based learning through PRAMIGO, BiomiX 3.0 enables researchers across disciplines with complementary tools to interrogate the biological sources of variation in their data, without requiring bioinformatics expertise.

bioinformatics↗