Search bioRxiv⌕ Search

Biology subjects

Adlard, D.

Publications and source records attributed to Adlard, D..

3 recordsLinked to original sources

Rapidly and reproducibly building a comprehensive catalogueof resistance-associated variants for M. tuberculosis

BackgroundCatalogues of genetic variants associated with resistance underpin whole-genome sequencing (WGS)-based predictions of drug susceptibility in Mycobacterium tuberculosis, and are essential for molecular diagnostics and surveillance. The current gold standard catalogues are those released by the WHO but the underlying data are not fully released and they are difficult to interpret. Open and reproducible methods would help address these problems, extending the important work already done. MethodsWe have developed an automated method, catomatic, that uses a binomial test to associate informative isolates with resistance or susceptibility, and built a catalogue (catomatic-1) from the same 39,358 samples used to construct the first edition of the WHO catalogue (WHOv1). We performed a sensitivity analysis to optimise statistical and bioinformatic parameters for each drug, and benchmarked catomatic-1 against WHOv1 using an independent Validation Dataset of 14,380 isolates. FindingsBy using simpler statistics, catomatic-1 algorithmically classified 1,329 genetic variants, ranging from five for linezolid to 440 for pyrazinamide. WHOv1 included generalisable rules added by a panel of experts, increasing its predictive coverage, but at the cost of reproducibility. Despite not including such expert rules, catomatic-1 achieves comparable performance for all drugs, with sensitivities for first-line agents above 88% on the independent Validation Dataset. The automated process allowed us to efficiently explore parameter space; for instance, detecting resistant variants with low read support improved the sensitivity for all drugs. InterpretationPerformant resistance catalogues for M. tuberculosis can be built automatically using transparent and reproducible statistical methods. As more data are collected, catalogue content and performance will evolve, highlighting the need for proper versioning, machine/human readability, and open access. This approach demonstrates resistance catalogues used in surveillance and diagnostics can be rapidly and reproducibily updated. FundingThe National Institute for Health and Care Research (NIHR), Engineering and Physics Sciences Research Council (EPSRC) and ORACLE Corporation. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed and preprint servers (bioRxiv, medRxiv), and publicly available mutation catalogues for studies linking Mycobacterium tuberculosis genomic variants with drug resistance using whole-genome or targeted sequencing and phenotypic drug-susceptibility testing (pDST). Search terms combined "Mycobacterium tuberculosis", "genome sequencing", "mutation catalogue", "mutation effects", "drug resistance", and individual drug names, with no language or date restriction. We included studies providing paired, clinical genomic and pDST or MIC data, excluding purely in-silico or case-only reports. This work directly builds on methodologies and data published by five prior studies, and makes primary comparisons with the First (WHOv1) and Second (WHOv2) Editions of the WHO Catalogue of mutations in Mycobacterium tuberculosis. Added value of this studyWe developed catomatic, a transparent, reproducible tool for building catalogues of resistance- and susceptibility-associated genetic variants. Trained on the same samples used to build WHOv1 and benchmarked on an independent Validation Dataset, catomatic achieves comparable sensitivity, specificity, and definitive prediction rates to WHOv1 without expert-rule augmentation and despite using simpler statistics. It optimises parameters per drug, produces machine-readable outputs (CSV/JSON), and demonstrates that adjusting read-support thresholds can improve detection of minor resistance subpopulations. Implications of all the available evidenceCatalogues of resistance-associated variants for M. tuberculosis can be rapidly and transparently constructed. Making catalogues available in human/machine-readable formats with uncertainty estimates will improve uptake of WGS for M. tuberculosis surveillance and diagnostics; using a reproducible process permits diagnostic test manufacturers, researchers, clinical and public health laboratories to select the level of statistical support necessitated by their specific use-case, Policymakers should balance the benefits of expert rules against loss of reproducibility. Future work will expand the size of the datasets used, integrate minimum inhibitory concentration data, and establish consensus workflows for routine, transparent catalogue updates.

microbiology↗

An improved catalogue for whole-genome sequencing prediction of bedaquiline resistance in M. tuberculosis using a reproduciblealgorithmic approach.

Bedaquiline (BDQ) has only been approved for use for a little over a decade yet is a key drug for treating multi-drug resistant tuberculosis, however rising levels of resistance threaten to reduce its effectiveness. Catalogues of mutations associated with resistance to bedaquiline are key to detecting resistance genetically for either diagnosis or surveillance. At present building catalogues requires considerable expert knowledge, often requires the use of complex grading rules, and is an irreproducible process. We developed an automated method, catomatic, that associates genetic variants with resistance (or susceptibility) using a two-tailed binomial test with a stated background rate and applied it to a dataset of 11,867 Mycobacterium tuberculosis samples with whole genome and bedaquline susceptibility testing data. Using this framework we investigated how to best classify variants and the phenotypic significance of minor alleles. The genes mmpS5 and mmpL5 are not directly associated with bedaquline resistance, and our catalogue of Rv0678, atpE, and pepQ variants attains a cross-validated sensitivity and specificity of 79.4 {+/-} 1.8 % and 98.5 {+/-} 0.3%, respectively, for 94 {+/-} 0.4% of samples. Identifying samples with subpopulations containing Rv0678 variants improves sensitivity, and detection thresholds in bioinformatic pipelines should therefore be lowered. By using a more permissive and deterministic algorithm trained on a sufficient number of resistant samples we have reproducibly constructed an AMR catalogue for BDQ resistance-associated variants that is comprehensive and accurate. Impact StatementBedaquiline has recently received global endorsement for tuberculosis treatment, yet the genetic determinants of antimicrobial resistance remain incompletely understood. Existing gold-standard methods for building mutation catalogs lack public accessibility and reproducibility. We introduce catomatic, a reproducible and publicly available method that employs simpler statistics to increase sensitivity to resistance-associated variants. This approach has enabled investigations into mechanisms of resistance, the significance of genetic subpopulations, and key data attributes that influence the ease of classifying effects and the accuracy of BDQ resistance phenotype prediction in clinical samples. We strongly emphasise the utility in using reproducible statistics and sustainably developed software in genetics-focussed microbiology.

microbiology↗

Predicting rifampicin resistance in M. tuberculosis using machine learning informed by protein structural and chemical features.

BackgroundRifampicin remains a key antibiotic in the treatment of tuberculosis. Despite advances in cataloguing resistance-associated variants (RAVs), novel and rare mutations in the relevent gene, rpoB, will be encountered in clinical samples, complicating the task of using genetics to predict whether a sample is resistant or not to rifampicin. We have trained a series of machine learning models with the aim of complementing genetics-based drug susceptibility testing. MethodsWe built a Test+Train dataset comprising 219 susceptible mutations and 46 RAVs. Features derived from the structure of the RNA polymerase or the change in chemistry introduced by the mutation were considered, however, only a few, notably the distance from the rifampicin binding site, were found to be predictive on their own. Due to the paucity of RAVs we used Monte Carlo cross-validation with 50 repeats to train four different machine learning models. ResultsAll four models behaved similarly with sensitivities and specificities in the range 0.84-0.88 and 0.94-0.97 although we preferred the ensemble of Decision Tree models as they are easy to inspect and understand. We showed that measuring distances from molecular dynamics simulations did not improve performance. ConclusionsIt is possible to predict whether a mutation in rpoB confers resistance to rifampicin using a machine learning model trained on a combination of structural, chemical and evolutionary features, however performance is moderate and training is complicated by the lack of data.

microbiology↗