Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.03.23.644786

Identifying Key Cells for Fibrosis by Systematically Calling Cell Type-Phenotype Associations across Massive Heterogenous Datasets

Abstract

Fibrotic diseases pose a significant burden on health care, yet the key pathogenic fibroblasts involved remain unclear. We developed the fibrotic disease fibroblast atlas (FDFA), which comprises 394 single-cell and 38 spatial transcriptomic samples from 11 common fibrotic diseases. To perform a cell-type phenotype association study in large-scale heterogeneous datasets, we developed the single-cell phenotype association research kit for large-scale dataset exploration (SPARKLE). SPARKLE handles heterogeneity by incorporating confounding information into its generalized linear mixed models (GLMMs). The application of SPARKLE to FDFA revealed that matrix fibroblasts (MTFs) constitute a crucial pathogenic cell group in fibrosis. Their increased proportion correlate with the fibrotic process. MTFs also synergize with MYO-Fs, increasing their degree of fibrosis. Based on MTF, we identified 25 potential antifibrotic targets for broad-spectrum antifibrotic therapies. This study enhances our understanding of fibrosis and provides a reliable framework for large-scale cell type-phenotype association research. HighlightsO_LIA cross-tissue, multidisease fibrotic disease fibroblast atlas (FDFA) reveals diverse fibroblast subtypes in fibrosis diseases. C_LIO_LIA new toolkit, SPARKLE, is introduced for detecting robust cell type-phenotype associations in large-scale heterogeneous datasets. C_LIO_LIMultiple pathogenic fibroblasts associated with common fibrosis disease (MTF) onset and progression have been identified, suggesting potential cell-specific therapeutic strategies for fibrosis-related diseases. C_LI

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chen, X., Ding, Y., Yang, S., Wei, L., Chu, M., Tang, H., Kong, L., Zhou, Y., Gao, G.. 2025-03-25. Identifying Key Cells for Fibrosis by Systematically Calling Cell Type-Phenotype Associations across Massive Heterogenous Datasets. https://doi.org/10.1101/2025.03.23.644786

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models

Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

bioinformatics↗

Cophylogeny simulators are not interchangeable: similarities, differences and structural biases in synthetic host-symbiont

Synthetic data are becoming increasingly important for computational studies of cophylogeny, including machine learning inference, benchmarking, and method testing. Several generators have been proposed to produce such data, but each relies on different assumptions about host-symbiont coevolution. These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used. In this article, we present a systematic structural analysis of representative cophylogeny generators under controlled scenarios. The goal is to make their assumptions explicit and to understand how these choices shape the synthetic data they produce as well as the conclusions that may be drawn from them.

bioinformatics↗

Evolutionary Diversification of Nitric Oxide Signaling Components Across Metazoa: A Comparative Phylogenomic Analysis

Nitric oxide (NO) is an evolutionarily ancient gaseous signaling molecule in animals, yet the evolutionary history of multiple components spanning NO synthesis, substrate regulation, sensing, and signal termination has not been examined in a single integrated phylogenetic framework across Metazoa. Here, we trace the phylogenetic and gene-tree/species-tree histories of ten core NO-pathway components across 89 eukaryotic proteomes spanning Amoebozoa, Excavata, Fungi, Archaeplastida, and Opisthokonta - including Porifera, Placozoa, Cnidaria, Ctenophora, and Bilateria - with a focus on the 69 opisthokont species that anchor the animal comparisons. The results reveal a strikingly modular evolutionary architecture. Nitric oxide synthase (NOS) is broadly conserved across bilaterian lineages, and reconciliation analyses show that NOS diversification was driven predominantly by speciation rather than lineage-specific duplication - supporting the relative conservation of NOS across the sampled bilaterian lineages . By contrast, the arginine-recycling enzymes ASS1 and ASL show ancestral duplications and inferred secondary losses in specific lineages, while arginase isoforms (ARG1/ARG2) and the cGMP-degrading enzyme PDE5A underwent extensive, independent duplications across metazoan groups. GUCY1A1- and GUCY1B1-like sequences were recovered across several metazoan lineages, but the two subunits exhibited partially divergent evolutionary trajectories. Together, these patterns support a model in which a comparatively conserved NO-producing component coexists with more dynamic diversification of associated pathway gene families, providing an evolutionary framework for investigating how NO-cGMP signaling may have been differentially deployed in nervous systems.

bioinformatics↗