Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.03.12.705671

SC-BIG: A Hierarchical Bayesian Model for Bulk-Informed Single Nucleotide Variant Calling in Single Cells

Abstract

Single-cell DNA sequencing (scDNA-seq) has emerged as a primary method for studying the evolution of cancer genomes and intra-tumor heterogeneity. However, despite technological advances, scDNA-seq remains noisy and is affected by amplification biases and allelic dropouts. Accurately determining the presence or absence of candidate somatic nucleotide variants (SNVs) in individual cancer cells therefore remains challenging. One strategy to alleviate this issue is to perform bulk whole-genome sequencing simultaneously with single-cell sequencing. To date, only few computational methods have been developed for bulk-informed detection of somatic SNVs in single cells, and existing methods do not adequately account for somatic copy-number alterations or clonal admixtures. We here present SC-BIG, a hierarchical Bayesian model that leverages bulk sequencing data from a representative tumor sample to improve SNV detection. SC-BIG propagates uncertainty across multiple biological parameters, including copy number alterations, sample purity, and SNV clonality. In a first step, the cancer cell fraction (CCF) of a SNV is jointly estimated from bulk and single-cell data. The CCF in turn then acts as a prior in the second inference step to calculate per-cell posterior probabilities for the presence of the SNV. We demonstrate that across simulated scenarios of varying CCFs, SC-BIG outperforms both naive thresholding and ProSolo, the only bulk-informed single-cell mutation caller described so far. Importantly, SC-BIG produces well-calibrated posterior probabilities that provide interpretable uncertainty quantification, enabling direct integration into downstream analyses such as phylogenetic reconstruction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Schuette, D., Kono, T. J. Y., Schwarz, R. F.. 2026-03-16. SC-BIG: A Hierarchical Bayesian Model for Bulk-Informed Single Nucleotide Variant Calling in Single Cells. https://doi.org/10.64898/2026.03.12.705671

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models

Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

bioinformatics↗

Cophylogeny simulators are not interchangeable: similarities, differences and structural biases in synthetic host-symbiont

Synthetic data are becoming increasingly important for computational studies of cophylogeny, including machine learning inference, benchmarking, and method testing. Several generators have been proposed to produce such data, but each relies on different assumptions about host-symbiont coevolution. These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used. In this article, we present a systematic structural analysis of representative cophylogeny generators under controlled scenarios. The goal is to make their assumptions explicit and to understand how these choices shape the synthetic data they produce as well as the conclusions that may be drawn from them.

bioinformatics↗

Evolutionary Diversification of Nitric Oxide Signaling Components Across Metazoa: A Comparative Phylogenomic Analysis

Nitric oxide (NO) is an evolutionarily ancient gaseous signaling molecule in animals, yet the evolutionary history of multiple components spanning NO synthesis, substrate regulation, sensing, and signal termination has not been examined in a single integrated phylogenetic framework across Metazoa. Here, we trace the phylogenetic and gene-tree/species-tree histories of ten core NO-pathway components across 89 eukaryotic proteomes spanning Amoebozoa, Excavata, Fungi, Archaeplastida, and Opisthokonta - including Porifera, Placozoa, Cnidaria, Ctenophora, and Bilateria - with a focus on the 69 opisthokont species that anchor the animal comparisons. The results reveal a strikingly modular evolutionary architecture. Nitric oxide synthase (NOS) is broadly conserved across bilaterian lineages, and reconciliation analyses show that NOS diversification was driven predominantly by speciation rather than lineage-specific duplication - supporting the relative conservation of NOS across the sampled bilaterian lineages . By contrast, the arginine-recycling enzymes ASS1 and ASL show ancestral duplications and inferred secondary losses in specific lineages, while arginase isoforms (ARG1/ARG2) and the cGMP-degrading enzyme PDE5A underwent extensive, independent duplications across metazoan groups. GUCY1A1- and GUCY1B1-like sequences were recovered across several metazoan lineages, but the two subunits exhibited partially divergent evolutionary trajectories. Together, these patterns support a model in which a comparatively conserved NO-producing component coexists with more dynamic diversification of associated pathway gene families, providing an evolutionary framework for investigating how NO-cGMP signaling may have been differentially deployed in nervous systems.

bioinformatics↗