Search bioRxivSearch

Biology subjects

Collins, F. S.

Publications and source records attributed to Collins, F. S..

3 recordsLinked to original sources

EndoC-βH1 multi-genomic profiling defines gene regulatory programs governing human pancreatic β cell identity and function

EndoC-{beta}H1 is emerging as a critical human beta cell model to study the genetic and environmental etiologies of beta cell function, especially in the context of diabetes. Comprehensive knowledge of its molecular landscape is lacking yet required to fully take advantage of this model. Here, we report extensive chromosomal (spectral karyotyping), genetic (genotyping), epigenetic (ChIP-seq, ATAC-seq), chromatin interaction (Hi-C, Pol2 ChIA-PET), and transcriptomic (RNA-seq, miRNA-seq) maps of this cell model. Integrated analyses of these maps define known (e.g., PDX1, ISL1) and putative (e.g., PCSK1, mir-375) beta cell-specific chromatin interactions and transcriptional cis-regulatory networks, and identify allelic effects on cis-regulatory element use and expression.\n\nImportantly, comparative analyses with maps generated in primary human islets/beta cells indicate substantial preservation of chromatin looping, but also highlight chromosomal heterogeneity and fetal genomic signatures in EndoC-{beta}H1. Together, these maps, and an interactive web application we have created for their exploration, provide important tools for the broad community in the design and success of experiments to probe and manipulate the genetic programs governing beta cell identity and (dys)function in diabetes.

genomics

BoostMe accurately predicts DNA methylation values in whole-genome bisulfite sequencing of multiple human tissues

BackgroundBisulfite sequencing is widely employed to study the role of DNA methylation in disease; however, the data suffer from biases due to variability in depth of coverage. Imputation of methylation values at low-coverage sites may mitigate these biases while also identifying important genomic features and motifs associated with predictive power.\n\nResultsHere we describe BoostMe, a novel method for imputation of DNA methylation within whole-genome bisulfite sequencing (WGBS) data based on a gradient boosting algorithm. Importantly, we designed a new feature that leverages information from multiple samples in the same tissue and disease state, enabling BoostMe to outperform existing imputation methods in speed and accuracy. We show that imputation improves WGBS concordance with the Infinium MethylationEPIC array at low WGBS sequencing depth, suggesting improvement in WGBS accuracy after imputation. Furthermore, we compare the ability of BoostMe and DeepCpG - a deep neural network method - to identify interesting features and motifs associated with methylation in three human tissues implicated in type 2 diabetes (T2D) etiology. We find that while BoostMe only identifies features important to general methylation levels across tissues, DeepCpG is able to learn differences in methylation-associated sequence motifs among different tissues and identify tissue-specific regulators of differentiation such as EBF1 in adipose, ASCL2 in muscle, and FOXA1, TCF12, and NRF1 in islets. Neither algorithm readily identified T2D-associated features.\n\nConclusionsOur findings demonstrate the current power and limitations of machine and deep learning algorithms to both improve the quality of, and infer biological meaning from, genome-wide DNA methylation data.

genomics

Interactions between genetic variation and cellular environment in skeletal muscle gene expression

From whole organisms to individual cells, responses to environmental conditions are influenced by genetic makeup, where the effect of genetic variation on a trait depends on the environmental context. RNA-sequencing quantifies gene expression as a molecular trait, and is capable of capturing both genetic and environmental effects. In this study, we explore opportunities of using allele-specific expression (ASE) to discover cis acting genotype-environment interactions (GxE) - genetic effects on gene expression that depend on an environmental condition. Treating 17 common, clinical traits as approximations of the cellular environment of 267 skeletal muscle biopsies, we identify 10 candidate interaction quantitative trait loci (iQTLs) across 6 traits (12 unique gene-environment trait pairs; 10% FDR per trait) including sex, systolic blood pressure, and low-density lipoprotein cholesterol. Although using ASE is in principle a promising approach to detect GxE effects, replication of such signals can be challenging as validation requires harmonization of environmental traits across cohorts and a sufficient sampling of heterozygotes for a transcribed SNP. Comprehensive discovery and replication will require large human transcriptome datasets, or the integration of multiple transcribed SNPs, coupled with standardized clinical phenotyping.

genetics