Search bioRxiv⌕ Search

Biology subjects

Odom-Mabey, A. R.

Publications and source records attributed to Odom-Mabey, A. R..

3 recordsLinked to original sources

Comparison of gene set scoring methods for reproducible evaluation of multiple tuberculosis gene signatures

RationaleMany blood-based transcriptional gene signatures for tuberculosis (TB) have been developed with potential use to diagnose disease, predict risk of progression from infection to disease, and monitor TB treatment outcomes. However, an unresolved issue is whether gene set enrichment analysis (GSEA) of the signature transcripts alone is sufficient for prediction and differentiation, or whether it is necessary to use the original statistical model created when the signature was derived. Intra-method comparison is complicated by the unavailability of original training data, missing details about the original trained model, and inadequate publicly-available software tools or source code implementing models. To facilitate these signatures replicability and appropriate utilization in TB research, comprehensive comparisons between gene set scoring methods with cross-data validation of original model implementations are needed. ObjectivesWe compared the performance of 19 TB gene signatures across 24 transcriptomic datasets using both re-rebuilt original models and gene set scoring methods to evaluate whether gene set scoring is a reasonable proxy to the performance of the original trained model. We have provided an open-access software implementation of the original models for all 19 signatures for future use. MethodsWe considered existing gene set scoring and machine learning methods, including ssGSEA, GSVA, PLAGE, Singscore, and Zscore, as alternative approaches to profile gene signature performance. The sample-size-weighted mean area under the curve (AUC) value was computed to measure each signatures performance across datasets. Correlation analysis and Wilcoxon paired tests were used to analyze the performance of enrichment methods with the original models. Measurement and Main ResultsFor many signatures, the predictions from gene set scoring methods were highly correlated and statistically equivalent to the results given by the original diagnostic models. PLAGE outperformed all other gene scoring methods. In some cases, PLAGE outperformed the original models when considering signatures weighted mean AUC values and the AUC results within individual studies. ConclusionGene set enrichment scoring of existing blood-based biomarker gene sets can distinguish patients with active TB disease from latent TB infection and other clinical conditions with equivalent or improved accuracy compared to the original methods and models. These data justify using gene set scoring methods of published TB gene signatures for predicting TB risk and treatment outcomes, especially when original models are difficult to apply or implement.

bioinformatics↗

Metagenomic profiling pipelines improve taxonomic classification for 16S amplicon sequencing data

Most experiments studying bacterial microbiomes rely on the PCR amplification of all or part of the gene for the 16S rRNA subunit, which serves as a biomarker for identifying and quantifying the various taxa present in a microbiome sample. Several computational methods exist for analyzing 16S amplicon-based metagenomics. However, the most-used bioinformatics tools cannot produce quality genus-level or species-level taxonomic calls and may underestimate the degree to which these calls are possible. We used 16S sequencing data from mock bacterial communities to evaluate the sensitivity and specificity of several bioinformatics pipelines and genomic reference libraries used for microbiome analyses, concentrating on measuring the accuracy of species-level taxonomic assignments of 16S amplicon reads. We evaluated the tools Qiime 2, Mothur, PathoScope 2, and Kraken 2 in conjunction with reference libraries from Greengenes, Silva, Kraken, and RefSeq. Profiling tools were compared using publicly available mock community data from several sources, comprising 136 samples with varied species richness and evenness, several different amplified regions within the 16S gene, and both DNA spike-ins and cDNA from collections of plated cells. PathoScope 2 and Kraken 2, both tools designed for whole-genome metagenomics, outperformed Qiime 2 using the DADA2 plugin and Mothur, both of which are theoretically specialized for 16S analyses.

bioinformatics↗

A comprehensive update to the Mycobacterium tuberculosis H37Rv reference genome

H37Rv is the most widely used M. tuberculosis strain. Its genome is globally used as the M. tuberculosis reference sequence. We developed Bact-Builder, a pipeline that leverages consensus building to generate complete and highly accurate gap-closed bacterial genomes and applied it to three independently sequenced cultures of a parental H37Rv laboratory stock. Two of the 4,417,942 base-pair long H37Rv assemblies were 100% identical, with the third differing by a single nucleotide. Compared to the existing H37Rv reference, the new sequence contained approximately 6.4 kb additional base pairs encoding ten new regions. These regions included insertions in PE/PPE genes and new paralogs of esxN and esxJ, which were differentially expressed compared to the reference genes. Additional sequencing and assembly with Bact-Builder confirmed that all 10 regions were also present in widely accepted strains of H37Rv: NR123 and TMC102. Bact-builder shows promise as an improved method to perform extremely accurate and reproducible de novo assemblies of bacterial genomes. Furthermore, our findings provide important updates to the primary tuberculosis reference genome.

microbiology↗