Search bioRxiv⌕ Search

Biology subjects

Yap, M.

Publications and source records attributed to Yap, M..

2 recordsLinked to original sources

Generalising uncertainty improves accuracy and safety of deep learning analytics applied to oncology

Trust and transparency are critical for deploying deep learning (DL) models into the clinic. DL application poses generalisation obstacles since training/development datasets often have different data distributions to clinical/production datasets that can lead to incorrect predictions with underestimated uncertainty. To investigate this pitfall, we benchmarked one pointwise and three approximate Bayesian DL models used to predict cancer of unknown primary with three independent RNA-seq datasets covering 10,968 samples across 57 primary cancer types. Our results highlight simple and scalable Bayesian DL significantly improves the generalisation of uncertainty estimation (e.g., p-value = 0.0013 for calibration). Moreover, we demonstrate Bayesian DL substantially improves accuracy under data distributional shifts when utilising uncertainty thresholding by designing a prototypical metric that evaluates the expected (accuracy) loss when deploying models from development to production, which we call the Area between Development and Production curve (ADP). In summary, Bayesian DL is a hopeful avenue of research for generalising uncertainty, which improves performance, transparency, and therefore safety of DL models for deployment in real-world.

bioinformatics↗

Microbiome-based environmental monitoring of a dairy processing facility highlights the challenges associated with low microbial-load samples

Food processing environments can harbor microorganisms responsible for food spoilage or foodborne disease. Efficient and accurate identification of microorganisms throughout the food chain can allow the identification of sources of contamination and the timely implementation of control measures. Currently, microbial monitoring of the food chain relies heavily on culture-based techniques. These assays are determined on the microbes expected to be present in the environment, and thus do not cater for unexpected contaminants. Many culture-based assays are also unable to distinguish between undesirable taxa and closely related harmless species. Furthermore, even when multiple culture-based approaches are used in parallel, it is still not possible to comprehensively characterize the entire microbiology of a food-chain sample. High throughput DNA sequencing represents a potential means through which microbial monitoring of the food chain can be enhanced. While sequencing platforms, such as the Illumina MiSeq, NextSeq and NovaSeq, are most typically found in research or commercial sequencing laboratories, newer portable platforms, such as the Oxford Nanopore Technologies (ONT) MinION, offer the potential for rapid analysis of food chain microbiomes. In this study, having initially assessed the ability of rapid MinION-based sequencing to discriminate between different microbes within a simple mock metagenomic mixture of related food spoilage, spore-forming microorganisms. Subsequently, we proceeded to compare the performance of both ONT and Illumina sequencing for environmental monitoring of an active food processing facility. Overall, ONT MinION sequencing provided accurate classification to species level, which was comparable to Illumina-derived outputs. However, while the MinION-based approach provided a means of easy library preparations and portability, the high concentrations of DNA needed to run the rapid sequencing protocols was a limiting factor, requiring the random amplification of template DNA in order to generate sufficient material for analysis.

microbiology↗