Search bioRxiv⌕ Search

Biology subjects

Flores, J. E.

Publications and source records attributed to Flores, J. E..

2 recordsLinked to original sources

Gaussian Mixture Modeling Extensions for Improved False Discovery Rate Estimation in GC-MS Metabolomics

The ability to reliably identify small molecules (e.g. metabolites) is key towards driving scientific advancement in metabolomics. Gas chromatography-mass spectrometry (GC-MS) is an analytic method that may be applied to facilitate this process. The typical GC-MS identification workflow involves quantifying the similarity of an observed sample spectrum and other features (e.g. retention index) to that of several references, noting the compound of the best-matching reference spectrum as the identified metabolite. While a deluge of similarity metrics exists, none quantify the error rate of generated identifications, thereby presenting an unknown risk of false identification or discovery. To quantify this unknown risk, we propose a model-based framework for estimating the false discovery rate (FDR) among a set of identifications. Extending a traditional mixture modeling framework, our method incorporates both similarity score and experimental information in estimating the FDR. We apply these models to identification lists derived from across 548 samples of varying complexity and sample type (e.g. fungal species, standard mixtures, etc.), comparing their performance to that of the traditional Gaussian mixture model (GMM). Through simulation, we additionally assess the impact of reference library size on the accuracy of FDR estimates. In comparing the best performing model extensions to the GMM, our results indicate relative decreases in median absolute estimation error (MAE) ranging from 12% to 70%, based on comparisons of the median MAEs across all hit-lists. Results indicate that these relative performance improvements generally hold despite library size, however FDR estimation error typically worsens as the set of reference compounds diminishes. For TOC graphic only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=74 SRC="FIGDIR/small/527348v1_ufig1.gif" ALT="Figure 1"> View larger version (11K): org.highwire.dtl.DTLVardef@ab9e2eorg.highwire.dtl.DTLVardef@11e0d8aorg.highwire.dtl.DTLVardef@ae2acorg.highwire.dtl.DTLVardef@a7ae20_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Evaluating Retention Index Score Assumptions to Refine GC-MS Metabolite Identification

As metabolomics grows into a high-throughput and high demand research field, current metrics for the identification of small molecules in GC-MS still require manual verification. Though steps have been taken to improve scoring metrics by combining spectral similarity (SS) and retention index (RI), the problem persists. A large body of literature has analyzed and refined SS scores, but few studies have explicitly studied improvements to RI scores. Here, we examined whether uninvestigated assumptions of the RI score are valid and propose ways to improve them. Query RI were matched to library RI with a generous window of +/-35 to avoid unintentional removal of valid compound identifications. Each match was manually verified as a true positive, true negative or unknown. Metabolites with at least 30 true positive identifications were included in downstream analyses, resulting in a total of 87 metabolites from samples of varying complexity and type (e.g., amino acid mixtures, human urine, fungal species, etc.). Our results showed that the RI score assumptions of normality, consistent variance across metabolites, and a mean error centered at 0 are often violated. We demonstrated through a cross-validation analysis that modifying these underlying assumptions according to empirical, metabolite-specific distributions improved the true positive and negative rankings. Further, we statistically determined the minimum number of samples required to estimate distributional parameters for scoring metrics. Overall, this work proposes a robust statistical pipeline to reduce the time bottleneck of metabolite identification by improving RI scores and thus minimizes the effort to complete manual verification.

bioinformatics↗