Search bioRxivSearch

Biology subjects

Benaglio, P.

Publications and source records attributed to Benaglio, P..

3 recordsLinked to original sources

Allele-specific NKX2-5 binding underlies multiple genetic associations with human EKG traits

Genetic variation affecting the binding of transcription factors (TFs) has been proposed as a major mechanism underlying susceptibility to common disease. NKX2-5, a key cardiac development TF, has been associated with electrocardiographic (EKG) traits through GWAS, but the extent to which differential binding of NKX2-5 contributes to these traits has not yet been studied. Here, we analyzed transcriptomic and epigenomic data generated from iPSC-derived cardiomyocyte lines (iPSC-CMs) from seven whole-genome sequenced individuals in a three-generational family. We identified ~2,000 single nucleotide variants (SNVs) associated with allele-specific effects (ASE) on NKX2-5 binding. These ASE-SNVs were enriched for altered TF motifs (both cognate and other cardiac TFs), and were positively correlated with changes in H3K27ac in iPSC-CMs, suggesting they impact cardiac enhancer activity. We found that NKX2-ASE-SNVs were significantly enriched for being heart-specific eQTLs and EKG GWAS variants, suggesting that altered NKX2-5 binding at multiple sites across the genome influences EKG traits. We used a fine-mapping approach to integrate iPSC-CM molecular phenotype data with a GWAS for heart rate, and determined that NKX2-5 ASE variants are likely causal for numerous known, as well as previously unidentified, heart rate loci. Analyzing Hi-C and gene expression data from iPSC-CMs at these heart rate loci, we identified several genes likely to be causally involved in heart rate variability. Our study demonstrates that differential binding of NKX2-5 is a common mechanism underlying genetic association with EKG traits, and shows that characterizing variants associated with differential binding of development TFs in iPSC-derived cell lines can identify novel loci and mechanisms influencing complex traits.

genetics

Integration of phased Hi-C and molecular phenotype data to study genetic and epigenetic effects on chromatin looping

While genetic variation at chromatin loops is relevant for human disease, the relationships between loop strength, genetics, gene expression, and epigenetics are unclear. Here, we quantitatively interrogate this relationship using Hi-C and molecular phenotype data across cell types and haplotypes. We find that chromatin loops consistently form across multiple cell types and quantitatively vary in strength, instead of exclusively forming within only one cell type. We show that large haplotype loop imbalance is primarily associated with imprinting and copy number variation, rather than genetically driven traits such as allele-specific expression. Finally, across cell types and haplotypes, we show that subtle changes in chromatin loop strength are associated with large differences in other molecular phenotypes, with a 2-fold change in looping corresponding to a 100-fold change in gene expression. Our study suggests that regulatory genetic variation could mediate its effects on gene expression through subtle modification of chromatin loop strength.

genomics

Using deep whole genome sequence, transcriptome and epigenome data to characterize the mutational burden of induced pluripotent stem cells

To understand the mutational burden of human induced pluripotent stem cells (iPSCs), we whole genome sequenced 18 fibroblast-derived iPSC lines and identified different classes of somatic mutations based on structure, origin and frequency. Copy number alterations affected 295 kb in each sample and strongly impacted gene expression. UV-damage mutations were present in ~45% of the iPSCs and accounted for most of the observed heterogeneity in mutation rates across lines. Subclonal mutations (not present in all iPSCs within a line) composed 10% of point mutations, and compared with clonal variants, showed an enrichment in active promoters and increased association with altered gene expression. Our study shows that, by combining WGS, transcriptome and epigenome data, we can understand the mutational burden of each iPSC line on an individual basis and suggests that this information could be used to prioritize iPSC lines for models of specific human diseases and/or transplantation therapy.

genomics