Search bioRxiv⌕ Search

Biology subjects

Yin, Y.-H.

Publications and source records attributed to Yin, Y.-H..

2 recordsLinked to original sources

CTDP: Identifying cell types associated with disease phenotypes using scRNA-seq data

Single-cell RNA sequencing enables transcriptome-wide analysis at single-cell resolution, offering unprecedented insights into cellular heterogeneity across biological conditions. However, accurately comparing transcriptomic distributions of cells from distinct biological states, such as healthy versus diseased individuals, remains challenging. To address this, we developed CTDP, a robust and interpretable computational framework that identifies disease phenotype-associated cell types of interest by integrating Lasso-regularized logistic regression with permutation testing. Through comprehensive evaluations on both simulated and real-world datasets, including melanoma immunotherapy, COVID-19 severity, and liver cirrhosis, CTDP consistently outperformed existing methods such as DA-seq, scDist, and PENCIL in both accuracy and robustness. In melanoma, CTDP uncovered immune-responsive clusters and revealed transcriptional regulators like PTPRC, CREM, and JUNB linked to immunotherapy efficacy. In COVID-19, it identified critical severity-associated cell types, such as B cells, NK cells, epithelial cells, and macrophages, which contribute to dysregulated immune responses and inflammation in severe cases. These results highlight CTDPs power in uncovering disease-relevant cell populations and its potential to advance precision medicine through single-cell analysis.

bioinformatics↗

Reference-guided genome assembly at scale using ultra-low-coverage high-fidelity long-reads with HiFiCCL

Population genomics using short-read resequencing captures single nucleotide polymorphisms and small insertions and deletions but struggles with structural variants (SVs), leading to a loss of heritability in genome-wide association studies. In recent years, long-read sequencing has improved pangenome construction for key eukaryotic species, addressing this issue to some extent. Sufficient-coverage high-fidelity (HiFi) data for population genomics is often prohibitively expensive, limiting its use in large-scale populations and broader eukaryotic species and creating an urgent need for robust ultra-low coverage assemblies. However, current assemblers underperform in such conditions. To address this, we propose HiFiCCL, the first assembly framework specifically designed for ultra-low-coverage high-fidelity reads, using a reference-guided, chromosome-by-chromosome assembly approach. We demonstrate that HiFiCCL improves ultra-low-coverage assembly performance of existing assemblers and outperforms the state-of-the-art assemblers on human and plant datasets. Tested on 45 human datasets ([~]5x coverage), HiFiCCL combined with hifiasm reduces the length of misassembled contigs relative to hifiasm by an average of 21.19% and up to 38.58%. These improved assemblies enhance germline structural variant detection, reduce chromosome-level mis-scaffolding, enable more accurate pangenome graph construction, and improve the detection of rare and somatic structural variants based on the pangenome graph under ultra-low-coverage conditions.

bioinformatics↗