Search bioRxivSearch

Biology subjects

Iqbal, Z.

Publications and source records attributed to Iqbal, Z..

6 recordsLinked to original sources

Real-time search of all bacterial and viral genomic data

Genome sequencing of pathogens is now ubiquitous in microbiology, and the sequence archives are effectively no longer searchable for arbitrary sequences. Furthermore, the exponential increase of these archives is likely to be further spurred by automated diagnostics. To unlock their use for scientific research and real-time surveillance we have combined knowledge about bacterial genetic variation with ideas used in web-search, to build a DNA search engine for microbial data that can grow incrementally. We indexed the complete global corpus of bacterial and viral whole genome sequence data (447,833 genomes), using four orders of magnitude less storage than previous methods. The method allows future scaling to millions of genomes. This renders the global archive accessible to sequence search, which we demonstrate with three applications: ultra-fast search for resistance genes MCR1-3, analysis of host-range for 2827 plasmids, and quantification of the rise of antibiotic resistance prevalence in the sequence archives.

bioinformatics

The global distribution and spread of the mobilized colistin resistance gene mcr-1

Colistin represents one of the very few available drugs for treating infections caused by carbapenem resistant Enterobacteriaceae (CRE). As such, the recent plasmid-mediated spread of the mobilized colistin resistance gene mcr-1 poses a significant public health threat requiring global monitoring and surveillance. In this work, we characterize the global distribution of mcr-1 using a dataset of 457 mcr-1 positive sequenced isolates consisting of currently publicly available mcr-1 carrying sequences combined with an additional 110 newly sequenced mcr-1 positive isolates from China. We find mcr-1 in a diversity of plasmid backgrounds but identify an immediate background common to all mcr-1 sequences. Our analyses establish that all mcr-1 elements in circulation descend from the same initial mobilization of mcr-1 by an ISApl1 transposon in the mid 2000s (2002-2008; 95% higher posterior density), followed by a dramatic demographic expansion, which led to its current global distribution. Our results provide the first systematic phylogenetic analysis of the origin and spread of mcr-1, and emphasize the importance of understanding the movement of mobile elements carrying antibiotic resistance genes across multiple levels of genomic organization.

microbiology

A novel multi SNP based method for identifying subspecies and associated lineages and sub-lineages of the Mycobacterium tuberculosis complex by whole genome sequencing

The clinical phenotype of zoonotic tuberculosis, its contribution to the global burden of disease and prevalence are poorly understood and probably underestimated. This is partly because currently available laboratory and in silico tools have not been calibrated to accurately identify all subspecies of the Mycobacterium tuberculosis complex (Mtbc). We here present the first such tool, SNPs to Identify TB ( SNP-IT). Applying SNP-IT to a collection of clinical genomes from a UK reference laboratory, we demonstrate an unexpectedly high number of M. orygis isolates. These are seen at a similar rate to M. bovis which attracts much health protection resource and yet M. orygis cases have not been previously described in the UK. From an international perspective it is possible that M. orygis is an underestimated zoonosis. As whole genome sequencing is increasingly integrated into the clinical setting, accurate subspecies identification with SNP-IT will allow the clinical phenotype, host range and transmission mechanisms of subspecies of the Mtbc to be studied in greater detail.

microbiology

Integrating long-range connectivity information into de Bruijn graphs

MotivationThe de Bruijn graph is a simple and efficient data structure that is used in many areas of sequence analysis including genome assembly, read error correction and variant calling. The data structure has a single parameter k, is straightforward to implement and is tractable for large genomes with high sequencing depth. It also enables representation of multiple samples simultaneously to facilitate comparison. However, unlike the string graph, a de Bruijn graph does not retain long range information that is inherent in the read data. For this reason, applications that rely on de Bruijn graphs can produce sub-optimal results given their input.\n\nResultsWe present a novel assembly graph data structure: the Linked de Bruijn Graph (LdBG). Constructed by adding annotations on top of a de Bruijn graph, it stores long range connectivity information through the graph. We show that with error-free data it is possible to losslessly store and recover sequence from a Linked de Bruijn graph. With assembly simulations we demonstrate that the LdBG data structure outperforms both the de Bruijn graph and the String Graph Assembler (SGA). Finally we apply the LdBG to Klebsiella pneumoniae short read data to make large (12 kbp) variant calls, which we validate using PacBio sequencing data, and to characterise the genomic context of drug-resistance genes.\n\nAvailabilityLinked de Bruijn Graphs and associated algorithms are implemented as part of McCortex, available under the MIT license at https://github.com/mcvean/mccortex.\n\nContactturner.isaac@gmail.com.

bioinformatics

Severe infections emerge from the microbiome by adaptive evolution

Bacteria responsible for the greatest global mortality colonize the human microbiome far more frequently than they cause severe infections. Whether mutation and selection within the microbiome accompany infection is unknown. We investigated de novo mutation in 1163 Staphylococcus aureus genomes from 105 infected patients with nose-colonization. We report that 72% of infections emerged from the microbiome, with infecting and nose-colonizing bacteria showing parallel adaptive differences. We found 2.8-to-3.6-fold enrichments of protein-altering variants in genes responding to rsp, which regulates surface antigens and toxicity; agr, which regulates quorum-sensing, toxicity and abscess formation; and host-derived antimicrobial peptides. Adaptive mutations in pathogenesis-associated genes were 3.1-fold enriched in infecting but not nose-colonizing bacteria. None of these signatures were observed in healthy carriers nor at the species-level, suggesting disease-associated, short-term, within-host selection pressures. Our results show that infection, like a cancer of the microbiome, emerges through spontaneous adaptive evolution, raising new possibilities for diagnosis and treatment.\n\nOne Sentence SummaryLife-threatening S. aureus infections emerge from nose microbiome bacteria in association with repeatable adaptive evolution.

genomics

Same-day diagnostic and surveillance data for tuberculosis via whole genome sequencing of direct respiratory samples.

Routine full characterization of Mycobacterium tuberculosis (TB) is culture-based, taking many weeks. Whole-genome sequencing (WGS) can generate antibiotic susceptibility profiles to inform treatment, augmented with strain information for global surveillance; such data could be transformative if provided at or near point of care.\n\nWe demonstrate a low-cost DNA extraction method for TB WGS direct from patient samples. We initially evaluated the method using the Illumina MiSeq sequencer (40 smear-positive respiratory samples, obtained after routine clinical testing, and 27 matched liquid cultures). M. tuberculosis was identified in all 39 samples from which DNA was successfully extracted. Sufficient data for antibiotic susceptibility prediction was obtained from 24 (62%) samples; all results were concordant with reference laboratory phenotypes. Phylogenetic placement was concordant between direct and cultured samples. Using an Illumina MiSeq/MiniSeq the workflow from patient sample to results can be completed in 44/16 hours at a cost of {pound}96/{pound}198 per sample.\n\nWe then employed a non-specific PCR-based library preparation method for sequencing on an Oxford Nanopore Technologies MinION sequencer. We applied this to cultured Mycobacterium bovis BCG strain (BCG), and to combined culture-negative sputum DNA and BCG DNA. For the latest flowcell, the estimated turnaround time from patient to identification of BCG was 6 hours, with full susceptibility and surveillance results 2 hours later. Antibiotic susceptibility predictions were fully concordant. A critical advantage of the MinION is the ability to continue sequencing until sufficient coverage is obtained, providing a potential solution to the problem of variable amounts of M. tuberculosis in direct samples.

microbiology