Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.08.17.608388

Low-cost and highly efficient generation of near-complete bacterial pathogen genomes by TELL-Seq

Abstract

Recent evidences suggest that de novo genome assembly can provide additional insights into genomic variation landscape beyond short-read NGS analyses. Despite the advent of longread sequencing technologies, generating high-quality bacterial genome assemblies remains expensive and requires high-quality and large quantities of DNA. TELL-Seq, an emerging linked-read technology, has been reported as a low-cost alternative for producing nearcomplete genome assembly in certain bacterial species. However, a systematic assessment of its performance and characteristics in de novo bacterial genome assembly has not yet been conducted. To address this, we benchmarked TELL-Seq using a set of clinical and standard bacterial pathogens with a wide range of genome size (2.0 6.4 Mbp), GC-content (32% - 68%), genome complexity (mappability from 0.15 to 0.60) and Gram types. Our findings indicate that, with 1/6 cost and 1/15,000 of DNA compared to PacBio HiFi, TELL-Seq could generate as high quality near-complete bacterial genome assemblies. The optimal depth for assembly is around 200x, as increased depth improves contiguity and completeness, though quality plateaus at 250x. In general, genome complexity was significantly correlated with assembly quality, but GC-content was not. Nevertheless, genomes with extreme GC-content may require higher depths for accurate assembly. Our study suggests that TELL-Seq could be a scalable method for large-scale bacterial genomic surveys. The data generated from this study could serve as a benchmark dataset for further algorithm development. Impact StatementAn increasing body of literature have shown that de novo genome assembles could provide additional insights into genomic variation beyond standard short-read technologies. Its particularly useful in clinical and epidemiological studies of bacteria, due to high variable nature of their genomes. Although great advances have been made in long-read sequencing technology, its still expensive with high requirement over sample DNAs quality and quantity, limiting its adoption in large-scale studies or surveys. Linked-read technology is an inexpensive alternative which used single molecule barcoding to mark origin of reads for better scaffolding of contigs assembled from normal short reads. TELL-Seq, as one of these promising technologies, was reported to be able to inexpensively and efficiently generate near-complete bacterial de novo genomic assemblies with low DNA inputs. Theres not yet a systematic benchmarking of this technology in bacterial pathogens. In this study, a set of clinically important bacterial pathogens with a wide spectrum of GC-content, genome size, genome complexity, and Gram Type were used to investigate the performance of TELL-Seq technology in de novo genome assembly. We found that, with about 1/6 of cost and 1/15000 of DNA quantity, it could generate comparable de novo assemblies compared to standard reference genomes or PacBio HiFi assemblies, while beating conventional short-read assemblies. We found several key factors, such as genome complexity, and sequencing depth, which would impact the quality of the assembly. Our study suggests that a sequencing depth of 200x is sufficient to achieve satisfactory results. Our study has revealed characteristics of this technology in bacterial genome assembly and paved the way for its large-scale application in clinical or epidemiological surveys. Data SummaryO_LIRaw sequencing data for all isolates have been deposited to SRA, accessible through these NCBI BioProject PRJNA1133244. C_LIO_LIA full list of SRA BioSample accession numbers are available in Tab S5. C_LIO_LIAssemblies and analyses companion this paper is available at: https://github.com/x-lab/tellseq_bacteria C_LI

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wu, Z., Zhang, T., Li, S., Shen, L., Yang, Z., Yang, Y., Chen, X., Li, B., Zhou, S., Zhou, X., Wu, B., Jiang, J., Li, X.. 2024-08-19. Low-cost and highly efficient generation of near-complete bacterial pathogen genomes by TELL-Seq. https://doi.org/10.1101/2024.08.17.608388

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A meta-interaction basis for cell-cell communication in tissues

Tissue function depends on signals exchanged between cells and the responses they elicit. Yet whether diverse cell-cell interactions in situ form recurrent sender-receiver programs remains unclear. We present SpiderNet, an interpretable representation-learning framework that discovers such directed programs as a compact basis of cell-cell meta-interactions (MIs) from spatial transcriptomics. SpiderNet jointly learns which sender regulators, ligand-receptor pairs, and receiver targets define each MI and where each program is active across neighboring cell pairs. The resulting representation traces multicellular relays and links communication to cell states, perturbation responses, and phenotypes. SpiderNet recovers ground-truth MIs and their molecular components in simulations and, in real tissues, shows stronger direction-specific agreement with independently curated regulatory programs in senders and receivers than alternative methods. Across more than 5.8 million spatially profiled cells, SpiderNet resolves an SPP1-THBS relay linking monocytes, fibroblasts, and tumor cells within an immune-suppressive ovarian cancer niche, predicts T-cell responses to held-out melanoma-cell perturbations, and identifies a T-cell-associated brain-aging program and age-predictive signals that transfer across regions and platforms. It reveals a recurrent pan-cancer COLLAGEN-linked fibroblast-tumor program whose projected abundance in independent cohorts is associated with poorer survival and non-response to immunotherapy. SpiderNet thus establishes MIs as a reusable organizational layer between molecular interactions and tissue phenotypes, providing a framework to resolve, compare, trace, and perturb multicellular regulation in situ.

bioinformatics↗

Heterogeneous Graph Contrastive Learning for Drug-Gene-Disease Motif Prediction

Drug repurposing and target discovery offer critical strategies for advancing therapeutic development by uncovering the potential biological pathways and novel associations among drugs, genes, and diseases. However, experimental discovery remains expensive and time-consuming, which limits the scalability of large-scale studies. In addition, existing computational approaches often struggle to effectively integrate heterogeneous biomedical data, capture the complex higher-order topological signatures of biological interactomes, and generalize to unseen entities. Here, we present HANAMI (Heterogeneous grAph coNtrastive leArning for drug-gene-disease Motif predIction), a multi-view deep graph learning framework designed to model complex interactions among drugs, genes, and diseases. HANAMI integrates diverse heterogeneous biomedical knowledge, including chemical structures, genomic sequences, and clinical phenotypes, and leverages relation-aware topology encoding, structure-aware aggregation, and contrastive learning to enable accurate motif prediction with biological context from the network. Systematic evaluation on benchmark datasets shows that HANAMI achieves up to 6% improvements over existing state-of-the-art methods in predicting drug-gene-disease motifs. The framework further demonstrates strong inductive generalization, maintaining an [~]18% performance advantage in zero-shot settings involving previously unseen entities. Beyond predictive performance, HANAMI effectively prioritizes drug-disease relationships investigated in Phase II or III trials while identifying candidate genes that suggest plausible mechanistic links. Together, HANAMI provides a computational framework for interpreting complex biomedical interactions, offering a scalable foundation to accelerate drug repurposing and therapeutic innovation.

bioinformatics↗

PTMExplorer: A Multi-Dimensional Integrative Visualization Platform for Protein Post-Translational Modification Function and Structure

Deciphering the functions of post-translational modifications (PTMs) is a critical bridge connecting large-scale modification proteomics data to mechanistic studies. However, most existing tools for visualizing PTM omics data are limited to site catalogs or single-dimensional feature displays. They lack the capability to simultaneously map user-derived differential modification sites onto multi-dimensional contexts, including protein three-dimensional (3D) structure, evolutionary conservation, functional sites, and disease associations. This limitation makes it difficult for researchers to rapidly assess the biological importance of candidate sites from among a vast number of differentially modified sites. Here, we present PTMExplorer, an interactive platform for the multi-dimensional visualization of protein PTMs. PTMExplorer comprises three core modules: PTM Inspector, built upon ProtVista, provides a multi-track, sequence-feature integrated view incorporating intrinsically disordered region (IDR) prediction (via flDPnn), surface accessibility calculation (via FreeSASA), and UniProt functional annotations; PTM 3D Locator, leveraging the Nightingale/Mol* engine, anchors modification sites onto AlphaFold/Protein Data Bank (PDB) 3D structures through residue mapping via PDBe-SIFTS; and PTM Overview, utilizing the R circlize package, presents a panoramic polar circos plot illustrating modification distribution and inter-group differential regulation. Additionally, three major disease-associated modification databases (PTMD, qPTM, and PhosCancer) are integrated as PTM-Disease Nexus, enabling co-localization comparison between user-defined differential sites and reported disease-related sites. PTMExplorer currently supports eight model organisms, accepts user-uploaded differential analysis results, and provides multi-dimensional annotations and various visualization options (https://www.bioladder.cn/PTMExplorer/). Using a multi-omics dataset from hepatocellular carcinoma (18 patients, 9 modification types) as a case study, we demonstrate the practical utility of PTMExplorer in screening potential biomarkers, revealing multi-modification coordination mechanisms, and distinguishing between absolute and relative quantification patterns.

bioinformatics↗