Search bioRxivSearch

Biology subjects

Mathema, B.

Publications and source records attributed to Mathema, B..

2 recordsLinked to original sources

Genotypic clustering does not imply recent tuberculosis transmission in a high prevalence setting: A genomic epidemiology study in Lima, Peru

BackgroundWhole genome sequencing (WGS) can elucidate Mycobacterium tuberculosis (Mtb) transmission patterns but more data is needed to guide its use in high-burden settings. In a household-based transmissibility study of 4,000 TB patients in Lima, Peru, we identified a large MIRU-VNTR Mtb cluster with a range of resistance phenotypes and studied host and bacterial factors contributing to its spread.\n\nMethodsWGS was performed on 61 of 148 isolates in the cluster. We compared transmission link inference using epidemiological or genomic data with and without the inclusion of controversial variants, and estimated the dates of emergence of the cluster and antimicrobial drug resistance acquisition events by generating a time-calibrated phylogeny. We validated our findings in genomic data from an outbreak of 325 TB cases in London. Using a larger set of 12,032 public Mtb genomes, we determined bacterial factors characterizing this cluster and under positive selection in other Mtb lineages.\n\nFindingsFour isolates were distantly related and the remaining 57 isolates diverged ca. 1968 (95% HPD: 1945-1985). Isoniazid resistance arose once, whereas rifampicin resistance emerged subsequently at least three times. Amplification of other drug resistance occurred as recently as within the last year of sampling. High quality PE/PPE variants and indels added information for transmission inference. We identified five cluster-defining SNPs, including esxV S23L to be potentially contributing to transmissibility.\n\nInterpretationClusters defined by MIRU-VNTR typing, could be circulating for decades in a high-burden setting. WGS allows for an improved understanding of transmission, as well as bacterial resistance and fitness factors.\n\nFundingThe study was funded by the National Institutes of Health (Peru Epi study U19-AI076217 and K01-ES026835 to MRF). The funding sources had no role in any aspect of the study, manuscript or decision to submit it for publication.\n\nResearch in contextO_ST_ABSEvidence before this studyC_ST_ABSUse of whole genome sequencing (WGS) to study tuberculosis (TB) transmission has proven to have higher resolution that traditional typing methods in low-burden settings. The implications of its use in high-burden settings are not well understood.\n\nAdded value of this studyUsing WGS, we found that TB clusters defined by traditional typing methods may be circulating for several decades. Genomic regions typically excluded from WGS analysis contain large amount of genetic variation that may affect interpretation of transmission events. We also identified five bacterial mutations that may contribute to transmission fitness.\n\nImplications of all the available evidenceAdded value of WGS for understanding TB transmission may be even higher in high-burden vs. low-burden settings. Methods integrating variants found in polymorphic sites and insertions and deletions are likely to have higher resolution. Several host and bacterial factors may be responsible for higher transmissibility that can be targets of intervention to interrupt TB transmission in communities.

microbiology

Beyond the SNP threshold: identifying outbreak clusters using inferred transmissions

Whole genome sequencing (WGS) is increasingly used to aid in understanding pathogen transmission [1]. Very often the number of single nucleotide polymorphisms (SNPs) separating isolates collected during an epidemiological study are used to identify sets of cases that are potentially linked by direct transmission. However, there is little agreement in the literature as to what an appropriate SNP cut-off threshold should be, or indeed whether a simple SNP threshold is appropriate for identifying sets of isolates to be treated as \"transmission clusters\". The SNP thresholds that have been adopted for inferring transmission vary widely even for one pathogen. As an alternative to reliance on a strict SNP threshold, we suggest that the key inferential target when studying the spread of an infectious disease is the number of transmission events separating cases. Here we describe a new framework for deciding whether two pathogen genomes should be considered as part of the same transmission cluster, based jointly on the number of SNP differences and the length of time over which those differences have accumulated. Our approach allows us to probabilistically characterize the number of inferred transmission events that separate cases. We show how this framework can be modified to consider variable mutation rates across the genome (e.g. SNPs associated with drug resistance) and we indicate how the methodology can be extended to incorporate epidemiological data such as spatial proximity. We use recent data collected from tuberculosis studies from British Columbia, Canada and the Republic of Moldova to apply and compare our clustering method to the SNP threshold approach. In the British Columbia data, different cases break off from the main clusters as cut-off thresholds are lowered; the transmission-based method obtains slightly different clusters than the SNP cut-offs. For the Moldova data, straightforward application of the methods shows no appreciable difference, but when we take into account the fact that resistance conferring sites likely do not follow the same mutation clock as most sites due to selection, the transmission-based approach differs from the SNP cut-off method. Outbreak simulations confirm that our transmission based method is at least as good at identifying direct transmissions as a SNP cut-off. We conclude that the new method is a promising step towards establishing a more robust identification of outbreaks.

genomics