Search bioRxiv⌕ Search

Biology subjects

Kundaje, A. B.

Publications and source records attributed to Kundaje, A. B..

2 recordsLinked to original sources

The ENCODE Imputation Challenge: A critical assessment of methods for cross-cell type imputation of epigenomic profiles

Functional genomics experiments are invaluable for understanding mechanisms of gene regulation. However, comprehensively performing all such experiments, even across a fixed set of sample and assay types, is often infeasible in practice. A promising alternative to performing experiments exhaustively is to, instead, perform a core set of experiments and subsequently use machine learning methods to impute the remaining experiments. However, questions remain as to the quality of the imputations, the best approaches for performing imputations, and even what performance measures meaningfully evaluate performance of such models. In this work, we address these questions by comprehensively analyzing imputations from 23 imputation models submitted to the ENCODE Imputation Challenge. We find that measuring the quality of imputations is significantly more challenging than reported in the literature, and is confounded by three factors: major distributional shifts that arise because of differences in data collection and processing over time, the amount of available data per cell type, and redundancy among performance measures. Our systematic analyses suggest several steps that are necessary, but also simple, for fairly evaluating the performance of such models, as well as promising directions for more robust research in this area.

bioinformatics↗

Cell-specific chromatin landscape of human coronary artery resolves regulatory mechanisms of disease risk

Coronary artery disease (CAD) is a complex inflammatory disease involving genetic influences across several cell types. Genome-wide association studies (GWAS) have identified over 170 loci associated with CAD, where the majority of risk variants reside in noncoding DNA sequences impacting cis-regulatory elements (CREs). Here, we applied single-cell ATAC-seq to profile 28,316 cells across coronary artery segments from 41 patients with varying stages of CAD, which revealed 14 distinct cellular clusters. We mapped ~320,000 accessible sites across all cells, identified cell type-specific elements, transcription factors, and prioritized functional CAD risk variants via quantitative trait locus and sequence-based predictive modeling. We identified a number of candidate mechanisms for smooth muscle cell transition states and identified putative binding sites for risk variants. We further employed CRE to gene linkage to nominate disease-associated key driver transcription factors such as PRDM16 and TBX2. This single cell atlas provides a critical step towards interpreting cis-regulatory mechanisms in the vessel wall across the continuum of CAD risk.

genomics↗