Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.09.26.754667

Accurate and scalable demultiplexing of single-cell RNA sequencing using BEACON

Abstract

Barcode-based multiplexing strategies can significantly increase sample throughput and decrease costs, while mitigating batch effects of single-cell RNA sequencing experiments. However, these approaches can be limited by inaccurate or inefficient demultiplexing, resulting in cell loss and reduced statistical power. Here, we present BEACON, a novel sample demultiplexing method that efficiently learns the background count distribution from data and removes it from individual cells, thereby improving classification accuracy. BEACON outperforms other state-of-the-art methods on multiple human data sets. We apply it to cancer cell line time course experiments in vitro, enabling the identification of genes associated with aggressive tumors in vivo. Finally, we adapt BEACON to multimodal protein-transcriptome profiling, enhancing protein signal recovery to identify a CD161-positive effector memory CD4 T-cell population with a Th17-like phenotype, which we prospectively validate. BEACON can therefore be applied to other droplet-based single-cell sequencing methodologies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ji, B. W., Lammers, M., Houser, A., Chan, N., Rodriguez, J., Chae, H., Xiang, C., Bui, J., Li, H., Ji, A.. 2026-09-30. Accurate and scalable demultiplexing of single-cell RNA sequencing using BEACON. https://doi.org/10.64898/2026.09.26.754667

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A moving target: non-stationary selection governs unsupervised prediction of viral fitness

Anticipating how mutations change viral fitness is central to genomic surveillance and vaccine design, yet the supervised phenotype data behind the most accurate variant-effect predictors are unavailable for most emerging pathogens. We ask how far label-free scoring can go using only sequences, their evolutionary history, and structure. We assemble a modular, fully unsupervised pipeline that estimates a few interpretable terms (intrinsic replicative fitness, antigenic escape, and realized growth), and that lets each term be produced by more than one estimator, so the estimator itself becomes a testable modeling choice. Benchmarking the intrinsic term on 21 viral deep-mutational-scanning assays from ProteinGym, we find that a 650-million-parameter single-sequence protein language model predicts viral mutational fitness weakly and heterogeneously (mean Spearman 0.15), whereas a trivial site-independent alignment model more than doubles it (0.39, better on 17 of 21 assays), with the largest gains on the antigenic surface proteins where the language model fails. Yet the ordering reverses across 186 non-viral ProteinGym assays, where the language model instead exceeds the alignment model, localizing the weakness to viral families under-represented in the model's training data. Alignment-conditioned language models (MSA Transformer, Tranception) recover this accuracy but do not clearly exceed the simple alignment, so the decisive feature is the family alignment, not model scale or architecture. Our central result is evolutionary. Using dated samples of SARS-CoV-2 spike and influenza H3N2 hemagglutinin, we show that the epoch of the alignment is itself a leading, virus-specific determinant of accuracy. This traces to non-stationary selection: the site-specific amino-acid preferences drift over time, abruptly for spike at the emergence of Omicron and gradually for H3N2 hemagglutinin. A phylogenetic mutation-selection estimator does not match the far cheaper alignment model, falling significantly below it on matched data. Unsupervised viral fitness prediction is, then, as much an evolutionary problem as a modeling one.

bioinformatics↗

Gene-family-dependent thermodynamic effects of oncogenic mutations on nucleosome-DNA binding stability: a comparative molecular dynamics study across 22 cancer hotspots

Oncogenic mutations are known to occur at non-random rates and specific hotspots, yet the structural and thermodynamic factors underlying these hotspots remain poorly understood. This study presents molecular dynamics simulations of 22 cancer hotspots across 10 oncogene families, simulated as histone-DNA complexes. Across 18 of 22 mutations, thermodynamic effects were predominantly localized to within 6 [A] of the DNA-histone interface (mean capture 104%), with van der Waals interactions driving the effect at the thermodynamic extremes, indicating that oncogenic mutations alter nucleosome binding through precise local contact changes rather than global structural rearrangements. Further, analysis revealed a gene-family correlated pattern in thermodynamic stability. RAS family mutations showed a consistent trend toward nucleosome stabilization (mean {Delta}{Delta}G = -4.82 kcal mol-1), while kinase domain mutations trended toward destabilization (mean {Delta}{Delta}G = +40.12 kcal mol-1), a difference reaching statistical significance in this exploratory analysis (Mann-Whitney U, p = 0.008). This interface-specific mechanism, combined with the gene-family-correlated thermodynamic pattern, provides a biophysical framework for understanding nucleosome-level contributions to cancer hotspot mutation biology.

bioinformatics↗

digestome: a licence-clean marker-gene panel for functional profiling of anaerobic digestion microbiomes

Functional profiling of anaerobic digestion microbiomes is routinely performed against KEGG or MetaCyc, both of which require a paid licence for commercial use. This blocks fee-for-service analysis for biogas operators, the setting where the results have the most immediate operational value. We present digestome, a curated marker-gene panel for anaerobic digestion built exclusively on sources that are free for commercial use (NCBIfam, public domain; Pfam, CC0; Rhea, CC BY), together with a scorer that reports pathway completeness, branch capability and, for hydrolysis, whether the enzyme is built for export. Benchmarked against 1,401 metagenome-assembled genomes from 134 anaerobic digesters, the panel assigned no methanogenesis route to any of 1,361 non-methanogens, and every acetoclastic call fell within Methanosarcina or Methanothrix. The result held for 3,043 species representatives from the Genome Taxonomy Database, spanning 203 phyla, which are built differently from binned metagenomes: the only background genomes given a route were two archaea carrying mcrA, and none of the 90 anaerobic methane and alkane oxidisers was given one. Specificity against non-methanogens does not show that a methanogen gets the right route: with a family-level check, 59 of 156 acetoclastic calls in that sample fell on methylotrophic genera that do not use acetate, and a lineage policy removed them, with 14 more left unconfirmed in unnamed genera. Testing whether a catalytic domain shares a polypeptide with an export module reduced the genomes credited with cellulolytic capacity from 597 to 50.

bioinformatics↗