Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.07.22.739971

Comparison of nuisance function construction strategies for double machine learning causal inference in single-cell transcriptomics: shared unsupervised deep learning does not require cross-fitting

Abstract

Inferring "whether a change in the expression of a given gene causally affects the disease state" from observational single-cell transcriptomic data is one of the central problems in single-cell biology. The difficulty lies in confounding: cell state, batch, cell cycle, and the co-expression of other genes may all simultaneously influence the target gene (treatment variable T) and the disease label (outcome variable Y), so that naive correlation analysis cannot distinguish causation from covariation. Double machine learning (DML), via orthogonal scores and cross-fitting, allows machine learning to estimate high-dimensional nuisance functions, thereby addressing the causal inference problem in high-dimensional data. Constrained by computational resources, this study takes a small-sample dataset with p{approx}n (2,120 cells, 1,999 background genes) as the experimental testbed and systematically compares three nuisance function construction strategies under this critical condition; strategies for the n>>p regime are then addressed by theoretical argument. Using systemic lupus erythematosus (SLE) peripheral blood memory B cells (GSE189050, 2,120 cells), we construct a three nuisance function construction strategies x (in-sample / cross-fitting) 2x3 factorial experiment and compare them in terms of resolution, biological plausibility, stability, and deconfounding ability in causal effect estimation. The three strategies are: (S1) direct linear nuisance regression on the high-dimensional background genes without dimensionality reduction; (S2) learning a shared low-dimensional representation with an unsupervised autoencoder, with the treatment and outcome residuals sharing that representation; (S3) fitting the treatment and outcome with two independent deep networks. The results yield a clear three-part pattern with mechanistic meaning: (1) both modes of S1 fail; (2) S2 achieves the optimum in-sample and requires no cross-fitting; (3) S3 is "rescued" under cross-fitting and becomes effective. We explain this pattern starting from the convergence rate condition of DML (Sec.5 Theoretical Foundations) and give the applicability boundaries of the two strategies: S2 (shared unsupervised deep learning) achieves estimation quality comparable to standard cross-fitted DML at a tiny fraction of the computational cost, making it a feasible scheme for large-scale screening; S3 (dual networks + cross-fitting), as the standard DML recipe, can serve as a broad-spectrum reference control for S2 results.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ye, W., Jiang, X., Shen, F.. 2026-07-26. Comparison of nuisance function construction strategies for double machine learning causal inference in single-cell transcriptomics: shared unsupervised deep learning does not require cross-fitting. https://doi.org/10.64898/2026.07.22.739971

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

The impacts of exogenous noise on stochastic disease dynamics

Much of the literature and intuition associated with mathematical epidemiology is driven by deterministic models, which are a reasonable assumption when the population size is large. Stochastic models, especially individual based models, are however considered vital when dealing with small population sizes, especially at times of invasion or extinction. The overwhelming majority of these models (both deterministic and stochastic) assume that the underlying parameters are fixed (or follow a regular seasonal pattern). Here, we consider an analytic framework for dealing with randomly varying parameters through the use of stochastic differential equations - thereby capturing the action of external noisy processes such as weather. In particular, we focus on when the transmission rate, {beta}, varies as the solution to a Cox-Ingersoll-Ross Model, such that {beta} is gamma distributed with autocorrelation. We consider the impact of this parameter variation on a stochastic version of the Susceptible-Infected-Recovered model, and for this 'double-stochastic' model show through simulation and analytical results that exogenous noise increases the impact of stochasticity, potentially leading to more early extinctions, wider variations in the number of cases at equilibrium, but that early growth rate can be faster or slower depending on the precise parameters.

systems biology↗

Site-resolved spatial and structural interactome of a human cell

The spatial and structural arrangement of proteins determine virtually every process in human cells. We combined gentle subcellular fractionation by differential ultracentrifugation with cross-linking mass spectrometry to systematically map this cellular proteome architecture with residue-level evidence, identifying 164,146 residue-to-residue links in HEK293 cells. These links capture spatial protein arrangement at a resolution sufficient to pinpoint protein localizations at sub-organelle level, determine protein orientations within cellular membranes, and identify inter-organelle contact sites. The residue-level information provides evidence for 18,074 direct protein-protein interactions (PPIs), which we integrate into AlphaFold-based pipelines to nominate PPI-mediating short linear motifs and generate assembly models of large protein complexes. Guided by these spatial and structural readouts, we discover new PPIs within the endomembrane system that regulate the compartmental localization of trafficking machinery. Leveraging a network topology-driven strategy, we augment our HEK293 dataset with PPI data from different cell lines and methods, expanding the spatial and structural interactome of human cells.

systems biology↗

Sensory variation and behavioural degeneracy: a framework for interpreting heterogeneity in the gut-brain axis

Gut microbiome differences are frequently interpreted as reflecting underlying biological differences between individuals. When outcomes are mediated by behaviour, however, this mapping may be fundamentally non-unique. This limits causal inference in gut-brain research, where microbiome differences in autism and depression are routinely attributed to intrinsic neurobiology despite highly variable, overlapping findings. I built a minimal agent-based model grounded in the known sensory variation across the autism spectrum. Dietary behaviour emerges from latent sensory traits, including sensory drive, predictability preference, and context sensitivity, through reinforcement learning and environmental interaction. This behaviour shapes gut microbiome composition. Behavioural variation organizes endogenously into a continuum of specialist, opportunist, and explorer strategies that maps onto the autism sensory spectrum. The system is fundamentally degenerate. Similar microbiome states arise from distinct behavioural pathways. Similar dietary patterns emerge from divergent latent traits. This many-to-one mapping reflects the structural interaction of behaviour, learning, and environmental variability, not stochasticity alone. Microbiome similarity therefore does not uniquely identify underlying cause. As such, the model provides a theoretical framework for interpreting heterogeneity in gut microbiome research, particularly in autism, and generalizes to any condition where behaviour mediates between neural processes and ecological outcomes.

systems biology↗