Search bioRxiv⌕ Search

Biology subjects

Theunissen, L.

Publications and source records attributed to Theunissen, L..

2 recordsLinked to original sources

Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomics

Automatic cell type annotation methods assign cell type labels to new, unlabeled datasets by leveraging relationships from a reference RNA-seq atlas. However, new datasets may include labels absent from the reference dataset or exhibit feature distributions that diverge from it. These scenarios can significantly affect the reliability of annotation predictions, a factor often overlooked in current automatic annotation methods. The field of out-of-distribution detection (OOD), primarily focused on computer vision, addresses the identification of instances that differ from the training distribution. Implementing OOD methods in the context of novel cell type annotation and data shift detection for single-cell transcriptomics may enhance annotation accuracy and trustworthiness. We evaluate 6 OOD detection methods: LogitNorm, MC dropout, Ensembles, Energy- based OOD, Deep NN and Posterior networks, for their annotation and OOD detection performance in both synthetical and real-life application settings. We show that OOD detection methods are able to accurately detect novel cell types and show their promise to detect severe data shifts on non-integrated datasets. Moreover, integration of the OOD datasets improves annotation performance, but interferes with OOD detection, diminishing novel cell type capabilities of the OOD methods.

bioinformatics↗

Uncertainty-aware single-cell annotation with a hierarchical reject option

AbstractAutomatic cell type annotation methods assign cell type labels to new datasets by extracting relationships from a reference RNA-seq dataset. However, due to the limited resolution of gene expression features, there is always uncertainty present in the label assignment. To enhance the reliability and robustness of annotation, most machine learning methods address this uncertainty by providing a full reject option, i.e. when the predicted confidence score of a cell type label falls below a user-defined threshold, no label is assigned and no prediction is made. As a better alternative, some methods deploy hierarchical models and consider a so-called partial rejection by returning internal nodes of the hierarchy as label assignment. However, because a detailed experimental analysis of various rejection approaches is missing in the literature, there is no consensus on best practices, superiority of certain methods, and potential drawbacks associated with rejection. We evaluate three annotation approaches (1) full rejection (2) partial rejection and (3) no rejection for both flat and hierarchical probabilistic classifiers. Our findings indicate that hierarchical classifiers are superior when rejection is applied, with partial rejection being the preferred rejection approach, as it preserves a significant amount of label information. For optimal rejection implementation, the rejection threshold should be determined through careful examination of a methods rejection behavior. Without rejection, flat and hierarchical annotation perform equally well, as long as the cell type hierarchy accurately captures transcriptomic relationships.

bioinformatics↗