Search bioRxiv⌕ Search

Biology subjects

Chen, D. G.

Publications and source records attributed to Chen, D. G..

4 recordsLinked to original sources

Single-cell Tree-based Model for Genomic-Disease Association

The rapid maturation of single-cell multi-omics technologies has enabled unprecedented resolution for mapping disease states and identifying disease-associated biomarkers. In practice, biomarkers are often discovered through differential detection that treat genomic features as independent contributors to phenotypes, while the combinatorial interactions that drive clinical outcomes remain a practical challenge. We present scanCT (single-cell analysis of Clinical Tree), a tree-based framework that identifies groups of genomic features associated with distinct disease phenotypes in a highly interpretable manner. scanCT uses an unbiased, model-based variable-selection procedure for data-driven split selection, which is important for handling the diverse distributional properties of single-cell data across modalities. The tree architecture captures feature interaction effects, and the association modeling enables adjustment for confounding factors. We apply scanCT to longitudinal single-cell multi-omics COVID-19 datasets spanning diverse clinical outcomes and multiple time points per patient. scanCT identifies phenotype-specific gene and protein markers while accounting for age and sex, and it reveals interpretable synergistic marker combinations that help explain differences in patient clinical phenotypes.

bioinformatics↗

Learning Human T Cell Behaviors through Generative AI Embeddings of T Cell Receptors

T cells interact with the world through T cell receptors (TCRs). The extent to which TCRs determine T cell behavior has not been comprehensively characterized. Our Tarpon model leverages advances in generative artificial intelligence to synthesize large-scale (>1M sequences) TCR atlases across human development and diseases into actionable insights. Tarpon creates: 1) bespoke sampling functions generating realistic Ag-specific TCRs, 2) embeddings revealing CD4+ and CD8+ single-positive TCR repertoires as distinct with divergent physiochemical properties, and 3) cross-dataset mappings of T cell states that validate fetal CD4+ versus CD8+ TCR differences in adults and find fetal type I innate T cells to map to MAIT and KIR+ adult CD8+ T cells which we verify via whole transcriptome analysis. Tarpon is a resource as a reference of TCRs across human physiological states and as a computational framework to create interpretable TCR embeddings, via physicochemical associations, that have broad implications for the field.

bioinformatics↗

APMAT analysis reveals the association between CD8 T cell receptors, cognate antigen, and T cell phenotype and persistence

Elucidating the relationships between a class I peptide antigen, a CD8 T cell receptor (TCR) specific to that antigen, and the T cell phenotype that emerges following antigen stimulation, remains a mostly unsolved problem, largely due to the lack of large data sets that can be mined to resolve such relationships. Here, we describe Antigen-TCR Pairing and Multiomic Analysis of T-cells (APMAT), an integrated experimental-computational framework designed for the high-throughput capture and analysis of CD8 T cells, with paired antigen, TCR sequence, and single-cell transcriptome. Starting with 951 putative antigens representing a comprehensive survey of the SARS-CoV-2 viral proteome, we utilize APMAT for the capture and single cell analysis of CD8 T cells from 62 HLA A*02:01 COVID-19 participants. We leverage this unique, comprehensive dataset to integrate with peptide antigen properties, TCR CDR3 sequences, and T cell phenotypes to show that distinct physicochemical features of the antigen-TCR pairs strongly associate with both T cell phenotype and T cell persistence. This analysis suggests that CD8+ T cell phenotype following antigen stimulation is at least partially deterministic, rather than the result of stochastic biological properties.

immunology↗

CITE-seq analysis reveals human cytomegalovirus and diabetes-associated adaptive NK cell alterations in cardiovascular disease.

Coronary artery disease (CAD) is a leading cause of mortality worldwide with Diabetes and human cyto-megalovirus (HCMV) infection as risk factors. CADs influence on human NK cells is not well characterized. CITE-seq analysis of a CAD cohort of 61 patients revealed distinctly higher NK cell SPON2 expression and lower IFNG expression in severe CAD patients. Interestingly, HCMV+ patients displayed lower SPON2 ex-pression while diabetes status reversed the HCMV effect. Diabetes led to diminished adaptive Fc{varepsilon}RI{gamma}-/low NK cell frequencies and was associated with a higher PBMC IL15/TGFB transcript ratio, while TGFB in-creased in severe CAD. SPON2 expression corresponded to changes in conventional vs. adaptive NK cell frequencies, and SPON2/IFNG ratio decreased in inflamed plaque tissue with an increased adaptive NK cell gene signature and was increased in severe CAD patients. Our results indicate that the SPON2/IFNG ra-tio and adaptive NK cell gene signature associated with stenosis severity or inflammation in CAD.

immunology↗