Search bioRxiv⌕ Search

Biology subjects

Conery, M.

Publications and source records attributed to Conery, M..

4 recordsLinked to original sources

Single-cell multiome analysis supports α-to-β transdifferentiation in human pancreas

Spontaneous transdifferentiation of pancreatic glucagon-producing alpha to insulin-secreting beta-cells has been observed in mouse but not in human islets1. Here, we analyzed the largest single-cell dataset of human islets to date, composed of 650,000 cells across 121 deceased organ donors, in search of transitional cell states. By integrating single-cell RNA-seq, single-nucleus ATAC-seq and single-nucleus multiome (joint RNA and ATAC profiling) datasets generated by the Human Pancreas Analysis Program (HPAP)2,3 we identified two previously undescribed cell populations (c11 and c13 cells), which together represent transitional states between alpha- and beta-cells. Some c11 cells are insulin-positive while others are glucagon positive, but none are double-positive. C11 cells repress alpha-cell identity genes and activate beta-cell specific genes. Moreover, the transcriptomic and epigenetic profiles of c11 and c13 cells indicate a transitioning phenotype driven by lineage-specific transcription factors. Genetic lineage tracing in primary human islet cells confirmed alpha-to-beta cell transdifferentiation. C11 and c13 cells exist in all islet samples regardless of disease statuses, with type 2 diabetic samples having significantly more transitioning cells than matched non-diabetic controls. The discovery of these transitional cell types suggests a possibility for future therapy - transdifferentiating alpha-cells to beta-cell through activation of the c11 gene program.

systems biology↗

Accelerating Genome- and Phenome-Wide Association Studies using GPUs - A case study using data from the Million Veteran Program

The expansion of biobanks has significantly propelled genomic discoveries yet the sheer scale of data within these repositories poses formidable computational hurdles, particularly in handling extensive matrix operations required by prevailing statistical frameworks. In this work, we introduce computational optimizations to the SAIGE (Scalable and Accurate Implementation of Generalized Mixed Model) algorithm, notably employing a GPU-based distributed computing approach to tackle these challenges. We applied these optimizations to conduct a large-scale genome-wide association study (GWAS) across 2,068 phenotypes derived from electronic health records of 635,969 diverse participants from the Veterans Affairs (VA) Million Veteran Program (MVP). Our strategies enabled scaling up the analysis to over 6,000 nodes on the Department of Energy (DOE) Oak Ridge Leadership Computing Facility (OLCF) Summit High-Performance Computer (HPC), resulting in a 20-fold acceleration compared to the baseline model. We also provide a Docker container with our optimizations that was successfully used on multiple cloud infrastructures on UK Biobank and All of Us datasets where we showed significant time and cost benefits over the baseline SAIGE model.

genetics↗

GWAS-informed data integration and non-coding CRISPRi screen illuminate genetic etiology of bone mineral density

Over 1,100 independent signals have been identified with genome-wide association studies (GWAS) for bone mineral density (BMD), a key risk factor for mortality-increasing fragility fractures; however, the effector gene(s) for most remain unknown. Informed by a variant-to-gene mapping strategy implicating 89 non-coding elements predicted to regulate osteoblast gene expression at BMD GWAS loci, we executed a single-cell CRISPRi screen in human fetal osteoblasts (hFOBs). The BMD relevance of hFOBs was supported by heritability enrichment from stratified LD-score regression involving 98 cell types grouped into 15 tissues. 23 genes showed perturbation in the screen, with four (ARID5B, CC2D1B, EIF4G2, and NCOA3) exhibiting consistent effects upon siRNA knockdown on three measures of osteoblast maturation and mineralization. Lastly, additional heritability enrichments, genetic correlations, and multi-trait fine-mapping revealed unexpectedly that many BMD GWAS signals are pleiotropic and likely mediate their effects via non-bone tissues. Extending our CRISPRi screening approach to these tissues could play a key role in fully elucidating the etiology of BMD.

genomics↗

Regularized sequence-context mutational trees capture variation in mutation rates across the human genome

Germline mutation is the mechanism by which genetic variation in a population is created. Inferences derived from mutation rate models are fundamental to many population genetics inference methods. Previous models have demonstrated that nucleotides flanking polymorphic sites - the local sequence context - explain variation in the probability that a site is polymorphic. However, limitations to these models exist as the size of the local sequence context window expands. These include a lack of robustness to data sparsity at typical sample sizes, lack of regularization to generate parsimonious models and lack of quantified uncertainty in estimated rates to facilitate comparison between models. To address these limitations, we developed Baymer, a regularized Bayesian hierarchical tree model that captures the heterogeneous effect of sequence contexts on polymorphism probabilities. Baymer implements an adaptive Metropolis-within-Gibbs Markov Chain Monte Carlo sampling scheme to estimate the posterior distributions of sequence-context based probabilities that a site is polymorphic. We show that Baymer accurately infers polymorphism probabilities and well-calibrated posterior distributions, robustly handles data sparsity, appropriately regularizes to return parsimonious models, and scales computationally at least up to 9-mer context windows. We demonstrate application of Baymer in three ways - first, identifying differences in polymorphism probabilities between continental populations in the 1000 Genomes Phase 3 dataset, second, in a sparse data setting to examine the use of polymorphism models as a proxy for de novo mutation probabilities as a function of variant age, sequence context window size, and demographic history, and third, comparing model concordance between different great ape species. We find a shared context-dependent mutation rate architecture underlying our models, enabling a transfer-learning inspired strategy for modeling germline mutations. In summary, Baymer is an accurate polymorphism probability estimation algorithm that automatically adapts to data sparsity at different sequence context levels, thereby making efficient use of the available data. AUTHOR SUMMARYMany biological questions rely on accurate estimates of where and how frequently mutations arise in populations. One factor that has been shown to predict the probability that a mutation occurs is the local DNA sequence surrounding a potential site for mutation. It has been shown that increasing the size of local DNA sequence immediately surrounding a site improves prediction of where, what type, and how frequently the site is mutated. However, current methods struggle to take full advantage of this trend as well as capturing how certain our estimates are, in practice. We have designed a model, implemented in software (named Baymer), that is able to use large windows of sequence context to accurately model mutation probabilities in a computationally efficient manner. We use Baymer to identify specific DNA sequences that have the biggest impacts on mutability and apply the model to find motifs that have potentially evolved mutability between different human populations. We also apply it to show that germline mutations observed as polymorphic sites in humans - those that have occurred in our recent evolutionary history - can model very young mutations (de novo mutations) as well as polymorphism observed in populations of closely related great ape species.

genomics↗