Search bioRxiv⌕ Search

Biology subjects

Modvig, S.

Publications and source records attributed to Modvig, S..

2 recordsLinked to original sources

CyStainer: A transformer-based variational autoencoder for robust marker imputation in high-parameter cytometry

High parameter cytometry is essential for clinical diagnostics through precise immune cell profiling, improved patient stratification, and monitoring, while also enhancing the understanding of cellular responses in disease and therapeutic contexts. The amount of cytometry data is growing fast, and with that, the need to merge different datasets for unified analysis. Here, we present CyStainer, a transformer-based variational autoencoder that demonstrates competitive or superior performance to existing methods on several key tasks related to marker prediction. As a key novelty, we demonstrate that CyStainer can impute markers without having a set of shared backbone markers. We performed several benchmarks using real-world FACS, CyTOF, InfinityFlow and CITE-seq datasets to show that CyStainer is a robust and flexible tool for panel merging, marker imputation, dataset integration and virtual staining of unseen samples. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=43 SRC="FIGDIR/small/735235v1_ufig1.gif" ALT="Figure 1"> View larger version (15K): org.highwire.dtl.DTLVardef@1e76266org.highwire.dtl.DTLVardef@1ed5303org.highwire.dtl.DTLVardef@1e4ff76org.highwire.dtl.DTLVardef@13faa48_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

A Reproducible and Extensible Benchmark of Supervised Cell Type Annotation Tools for Cytometry Data

High-dimensional cytometry technologies such as flow cytometry (FCM) and mass cytometry (CyTOF) are central to immunophenotyping in research and clinical practice. While manual gating remains the standard for cell population annotation, it is time-consuming, difficult to scale, and subject to inter-operator variability. Supervised annotation methods have emerged as a way of scaling manual annotation work, yet independent benchmarks for comparing these tools remain limited and quickly become outdated. This study presents a reproducible and extensible benchmark of supervised cytometry annotation tools implemented within the OmniBenchmark framework. Five supervised annotation methods were evaluated, spanning linear models, nearest-neighbor approaches, tree-based classifiers, mixture-rule systems, and deep learning, across eight publicly available datasets carefully selected to cover technologies, tissues, panel designs, and healthy and disease contexts. Using a sample-centric cross-validation design that reflects common reference-mapping scenarios, overall and per-population F1 scores, performance on rare populations, runtime, and robustness to reduced training set sizes was tested. Performance varied substantially across datasets and was not fully explained by dataset size or dimensionality, highlighting both operator dependence in annotation and the importance of biological context, cohort heterogeneity, and population imbalance. Less prevalent populations (<1%) remained a key challenge for most methods. Downsampling analyses showed that moderate reference sizes were often sufficient to achieve near-maximum performance. Rather than ranking methods, this benchmark provides a standardized and transparent framework for evaluating annotation tools under realistic deployment conditions. As a living resource, the OmniBenchmark implementation supports continuous integration of new datasets, tools, and metrics for both tool developers and end users annotating datasets. This enables ongoing, reproducible method comparison and informed tool selection for diverse cytometry applications.

bioinformatics↗