Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.04.17.649405

Generative prediction of causal gene sets responsible for complex traits

Abstract

The relationship between genotype and phenotype remains an outstanding question for organism-level traits because these traits are generally complex. The challenge arises from complex traits being determined by a combination of multiple genes (or loci), which leads to an explosion of possible genotype-phenotype mappings. The primary techniques to resolve these mappings are genome/transcriptome-wide association studies, which are limited by their lack of causal inference and statistical power. Here, we develop an approach that leverages transcriptional data endowed with causal information and a generative machine learning model to strengthen statistical power. Our implementation of the approach--dubbed TWAVE--includes a variational autoencoder trained on human transcriptional data, which is incorporated into an optimization framework. Given a trait phenotype, TWAVE generates expression profiles, which we dimensionally reduce by identifying independently varying generalized pathways (eigengenes). We then conduct constrained optimization to find causal gene sets that are the gene perturbations whose measured transcriptomic responses best explain trait phenotype differences. By considering several complex traits, we show that the approach identifies causal genes that cannot be detected by the primary existing techniques. Moreover, the approach identifies complex diseases caused by distinct sets of genes, meaning that the disease is polygenic and exhibits distinct subtypes driven by different genotype-phenotype mappings. We suggest that the approach will enable the design of tailored experiments to identify multi-genic targets to address complex diseases. Significance summaryResearchers have long sought to bridge the gap between phenotypes and the genotypes that cause them. This gap remains open because current methods focus on associating phenotypes to a combinatorially explosive number of genotypic possibilities, resulting in a loss of statistical power. We overcome this limitation by employing transcriptomic data from complex, polygenic, human diseases combined with measured transcriptomic responses to gene perturbations in cell lines. The former data allow us to perform generative modeling and dimensional reduction to map transcriptome to phenotype, while the latter incorporate causal information regarding how gene regulation shapes phenotype. We predict sets of genes that explain the emergence of complex traits, which suggest possible multi-target disease treatments.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kuznets-Speck, B., Ogonor, B., Wytock, T. P., Motter, A. E.. 2025-04-22. Generative prediction of causal gene sets responsible for complex traits. https://doi.org/10.1101/2025.04.17.649405

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models

Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

bioinformatics↗

Cophylogeny simulators are not interchangeable: similarities, differences and structural biases in synthetic host-symbiont

Synthetic data are becoming increasingly important for computational studies of cophylogeny, including machine learning inference, benchmarking, and method testing. Several generators have been proposed to produce such data, but each relies on different assumptions about host-symbiont coevolution. These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used. In this article, we present a systematic structural analysis of representative cophylogeny generators under controlled scenarios. The goal is to make their assumptions explicit and to understand how these choices shape the synthetic data they produce as well as the conclusions that may be drawn from them.

bioinformatics↗

Evolutionary Diversification of Nitric Oxide Signaling Components Across Metazoa: A Comparative Phylogenomic Analysis

Nitric oxide (NO) is an evolutionarily ancient gaseous signaling molecule in animals, yet the evolutionary history of multiple components spanning NO synthesis, substrate regulation, sensing, and signal termination has not been examined in a single integrated phylogenetic framework across Metazoa. Here, we trace the phylogenetic and gene-tree/species-tree histories of ten core NO-pathway components across 89 eukaryotic proteomes spanning Amoebozoa, Excavata, Fungi, Archaeplastida, and Opisthokonta - including Porifera, Placozoa, Cnidaria, Ctenophora, and Bilateria - with a focus on the 69 opisthokont species that anchor the animal comparisons. The results reveal a strikingly modular evolutionary architecture. Nitric oxide synthase (NOS) is broadly conserved across bilaterian lineages, and reconciliation analyses show that NOS diversification was driven predominantly by speciation rather than lineage-specific duplication - supporting the relative conservation of NOS across the sampled bilaterian lineages . By contrast, the arginine-recycling enzymes ASS1 and ASL show ancestral duplications and inferred secondary losses in specific lineages, while arginase isoforms (ARG1/ARG2) and the cGMP-degrading enzyme PDE5A underwent extensive, independent duplications across metazoan groups. GUCY1A1- and GUCY1B1-like sequences were recovered across several metazoan lineages, but the two subunits exhibited partially divergent evolutionary trajectories. Together, these patterns support a model in which a comparatively conserved NO-producing component coexists with more dynamic diversification of associated pathway gene families, providing an evolutionary framework for investigating how NO-cGMP signaling may have been differentially deployed in nervous systems.

bioinformatics↗