Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomics
Automatic cell type annotation methods assign cell type labels to new, unlabeled datasets by leveraging relationships from a reference RNA-seq atlas. However, new datasets may include labels absent from the reference dataset or exhibit feature distributions that diverge from it. These scenarios can significantly affect the reliability of annotation predictions, a factor often overlooked in current automatic annotation methods. The field of out-of-distribution detection (OOD), primarily focused on computer vision, addresses the identification of instances that differ from the training distribution. Implementing OOD methods in the context of novel cell type annotation and data shift detection for single-cell transcriptomics may enhance annotation accuracy and trustworthiness. We evaluate 6 OOD detection methods: LogitNorm, MC dropout, Ensembles, Energy- based OOD, Deep NN and Posterior networks, for their annotation and OOD detection performance in both synthetical and real-life application settings. We show that OOD detection methods are able to accurately detect novel cell types and show their promise to detect severe data shifts on non-integrated datasets. Moreover, integration of the OOD datasets improves annotation performance, but interferes with OOD detection, diminishing novel cell type capabilities of the OOD methods.