Search bioRxivSearch

Biology subjects

Sankararaman, S.

Publications and source records attributed to Sankararaman, S..

10 recordsLinked to original sources

Cell-type-specific resolution epigenetics without the need for cell sorting or single-cell biology

High costs and technical limitations of cell sorting and single-cell techniques currently restrict the collection of large-scale, cell-type-specific DNA methylation data. This, in turn, impedes our ability to tackle key biological questions that pertain to variation within a population, such as identification of disease-associated genes at a cell-type-specific resolution. Here, we show mathematically and empirically that cell-type-specific methylation levels of an individual can be learned from its tissue-level bulk data, conceptually emulating the case where the individual has been profiled with a single-cell resolution and then signals were aggregated in each cell population separately. Provided with this unprecedented way to perform powerful large-scale epigenetic studies with cell-type-specific resolution, we revisit previous studies with tissue-level bulk methylation and reveal novel associations with leukocyte composition in blood and with rheumatoid arthritis. For the latter, we further show consistency with validation data collected from sorted leukocyte sub-types. Corresponding software is available from: https://github.com/cozygene/TCA.

bioinformatics

Preoperative predictions of in-hospital mortality using electronic medical record data

BackgroundPredicting preoperative in-hospital mortality using readily-available electronic medical record (EMR) data can aid clinicians in accurately and rapidly determining surgical risk. While previous work has shown that the American Society of Anesthesiologists (ASA) Physical Status Classification is a useful, though subjective, feature for predicting surgical outcomes, obtaining this classification requires a clinician to review the patients medical records. Our goal here is to create an improved risk score using electronic medical records and demonstrate its utility in predicting in-hospital mortality without requiring clinician-derived ASA scores.\n\nMethodsData from 49,513 surgical patients were used to train logistic regression, random forest, and gradient boosted tree classifiers for predicting in-hospital mortality. The features used are readily available before surgery from EMR databases. A gradient boosted tree regression model was trained to impute the ASA Physical Status Classification, and this new, imputed score was included as an additional feature to preoperatively predict in-hospital post-surgical mortality. The preoperative risk prediction was then used as an input feature to a deep neural network (DNN), along with intraoperative features, to predict postoperative in-hospital mortality risk. Performance was measured using the area under the receiver operating characteristic (ROC) curve (AUC).\n\nResultsWe found that the random forest classifier (AUC 0.921, 95%CI 0.908-0.934) outperforms logistic regression (AUC 0.871, 95%CI 0.841-0.900) and gradient boosted trees (AUC 0.897, 95%CI 0.881-0.912) in predicting in-hospital post-surgical mortality. Using logistic regression, the ASA Physical Status Classification score alone had an AUC of 0.865 (95%CI 0.848-0.882). Adding preoperative features to the ASA Physical Status Classification improved the random forest AUC to 0.929 (95%CI 0.915-0.943). Using only automatically obtained preoperative features with no clinician intervention, we found that the random forest model achieved an AUC of 0.921 (95%CI 0.908-0.934). Integrating the preoperative risk prediction into the DNN for postoperative risk prediction results in an AUC of 0.924 (95%CI 0.905-0.941), and with both a preoperative and postoperative risk score for each patient, we were able to show that the mortality risk changes over time.\n\nConclusionsFeatures easily extracted from EMR data can be used to preoperatively predict the risk of in-hospital post-surgical mortality in a fully automated fashion, with accuracy comparable to models trained on features that require clinical expertise. This preoperative risk score can then be compared to the postoperative risk score to show that the risk changes, and therefore should be monitored longitudinally over time.\n\nAuthor summaryRapid, preoperative identification of those patients at highest risk for medical complications is necessary to ensure that limited infrastructure and human resources are directed towards those most likely to benefit. Existing risk scores either lack specificity at the patient level, or utilize the American Society of Anesthesiologists (ASA) physical status classification, which requires a clinician to review the chart. In this manuscript we report on using machine-learning algorithms, specifically random forest, to create a fully automated score that predicts preoperative in-hospital mortality based solely on structured data available at the time of surgery. This score has a higher AUC than both the ASA physical status score and the Charlson comorbidity score. Additionally, we integrate this score with a previously published postoperative score to demonstrate the extent to which patient risk changes during the perioperative period.

bioinformatics

Haplotype-based eQTL mapping finds evidence for complex gene regulatory regions poorly tagged by marginal SNPs

MotivationExpression quantitative trait loci (eQTLs), variations in the genome that impact gene expression, are identified through eQTL studies that test for a relationship between single nucleotide polymorphisms (SNPs) and gene expression levels. These studies typically assume an underlying additive model. Non-additive tests have been proposed, but are limited due to the increase in the multiple testing burden and are potentially biased by filtering criteria that relies on marginal association data. Here we propose using combinations of short haplotypes instead of SNPs as predictors for gene expression. Essentially, this method looks for genomic regions where haplotypes have different effect sizes. The differences in effect can be due to multiple genetic architectures such as a single SNP, a burden of rare SNPs, multiple SNPs with independent effect or multiple SNPs with an interaction effect occurring on the same haplotype.\n\nResultsSimulations show that when haplotypes, rather than SNPs, are assigned non-zero effect sizes, our method has increased power compared to the marginal SNP method. In the GEUVADIS gene expression data, our method finds 101 more eGenes than the marginal method (5,202 vs. 5,101). The methods do not have full overlap in the eGenes that they find. Of the 5,202 eGenes found by our method, 707 are not found by the marginal method--even though it has a lower significance threshold. This indicates that many genes have regulatory architectures that are not well tagged by marginal SNPs and demonstrates the need to better model alternative archi-tectures.

genetics

Single cell RNAseq uncovers a robust transcriptional response to morphine by oligodendrocytes

Molecular and behavioral responses to opioids are thought to be primarily mediated by neurons, although there is accumulating evidence that other cell types also play a role in drug addiction. To investigate cell-type-specific opioid responses, we performed single-cell RNA sequencing of the nucleus accumbens of mice following acute morphine treatment. Differential expression analysis uncovered robust morphine-dependent changes in gene expression in oligodendrocytes. We examined the expression of selected genes, including Cdkn1a and Sgk1, by FISH, confirming their induction by morphine in oligodendrocytes. Further analysis using RNAseq of FACS-purified oligodendrocytes revealed a large cohort of morphine-regulated genes. Importantly, the affected genes are enriched for roles in cellular pathways intimately linked to oligodendrocyte maturation and myelination, including the unfolded protein response. Altogether, our data shed light on a novel, morphine-dependent transcriptional response by oligodendrocytes that may contribute to the myelination defects observed in human opioid addicts.

neuroscience

A unifying framework for joint trait analysis under a non-infinitesimal model

MotivationA large proportion of risk regions identified by genome-wide association studies (GWAS) are shared across multiple diseases and traits. Understanding whether this clustering is due to sharing of causal variants or chance colocalization can provide insights into shared etiology of complex traits and diseases.\n\nResultsIn this work, we propose a flexible, unifying framework to quantify the overlap between a pair of traits called UNITY (Unifying Non-Infinitesimal Trait analYsis). We formulate a Bayesian generative model that relates the overlap between pairs of traits to GWAS summary statistic data under a non-infinitesimal genetic architecture underlying each trait. We propose a Metropolis-Hastings sampler to compute the posterior density of the genetic overlap parameters in this model. We validate our method through comprehensive simulations and analyze summary statistics from height and BMI GWAS to show that it produces estimates consistent with the known genetic makeup of both traits.\n\nAvailabilityThe UNITY software is made freely available to the research community at: https://github.com/bogdanlab/UNITY\n\nContactruthjohnson@ucla.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

A scalable estimator of SNP heritability for Biobank-scale data

MotivationHeritability, the proportion of variation in a trait that can be explained by genetic variation, is an important parameter in efforts to understand the genetic architecture of complex phenotypes as well as in the design and interpretation of genome-wide association studies. Attempts to understand the heritability of complex phenotypes attributable to genome-wide SNP variation data has motivated the analysis of large datasets as well as the development of sophisticated tools to estimate heritability in these datasets.\n\nLinear Mixed Models (LMMs) have emerged as a key tool for heritability estimation where the parameters of the LMMs, i.e., the variance components, are related to the heritability attributable to the SNPs analyzed. Likelihood-based inference in LMMs, however, poses serious computational burdens.\n\nResultsWe propose a scalable randomized algorithm for estimating variance components in LMMs. Our method is based on a MoM estimator that has a runtime complexity [Formula] for N individuals and M SNPs (where B is a parameter that controls the number of random matrix-vector multiplications). Further, by leveraging the structure of the genotype matrix, we can reduce the time complexity to [Formula].\n\nWe demonstrate the scalability and accuracy of our method on simulated as well as on empirical data. On standard hardware, our method computes heritability on a dataset of 500, 000 individuals and 100, 000 SNPs in 38 minutes.\n\nAvailabilityThe RHE-reg software is made freely available to the research community at: https://github.com/sriramlab/RHE-reg\n\nContactsriram@cs.ucla.edu

bioinformatics

A unifying framework for summary statistic imputation

Imputation has been widely utilized to aid and interpret the results of Genome-Wide Association Studies(GWAS). Imputation can increase the power to identify associations when the causal variant was not directly observed or typed in the GWAS. There are two broad classes of methods for imputation. The first class imputes the genotypes at the untyped variants given the genotypes at the typed variants and then performs a statistical test of association at the imputed variants. The second class of methods, summary statistic imputation, directly imputes the association statics at the untyped variants given the association statistics observed at the typed variants. This second class of methods is appealing as it tends to be computationally efficient while only requiring the summary statistics from a study while the former class requires access to individual-level data that can be difficult to obtain. The statistical properties of these two classes of imputation methods have not been fully understood. In this paper, we show that the two classes of imputation methods are equivalent, i.e., have identical asymptotic multivariate normal distributions with zero mean and minor variations in the covariance matrix, under some reasonable assumptions. Using this equivalence, we can understand the effect of imputation methods on power. We show that a commonly employed modification of summary statistic imputation that we term summary statistic imputation with variance re-weighting generally leads to a loss in power. On the other hand, our proposed method, summary statistic imputation without performing variance re-weighting, fully accounts for imputation uncertainty while achieving better power.

bioinformatics

Recovering signals of ghost archaic admixture in the genomes of present-day Africans

While introgression from Neanderthals and Denisovans has been well-documented in modern humans outside Africa, the contribution of archaic hominins to the genetic variation of present-day Africans remains poorly understood. Using 405 whole-genome sequences from four sub-Saharan African populations, we provide complementary lines of evidence for archaic introgression into these populations. Our analyses of site frequency spectra indicate that these populations derive 2-19% of their genetic ancestry from an archaic population that diverged prior to the split of Neanderthals and modern humans. Using a method that can identify segments of archaic ancestry without the need for reference archaic genomes, we built genome-wide maps of archaic ancestry in the Yoruba and the Mende populations that recover about 482 and 502 megabases of archaic sequence, respectively. Analyses of these maps reveal segments of archaic ancestry at high frequency in these populations that represent potential targets of adaptive introgression. Our results reveal the substantial contribution of archaic ancestry in shaping the gene pool of present-day African populations. One sentence summaryMultiple present-day African populations inherited genes from an unknown archaic population that diverged before modern humans and Neanderthals split.

genomics

Natural selection interacts with the local recombination rate to shape the evolution of hybrid genomes

While hybridization between species is increasingly appreciated to be a common occurrence, little is known about the forces that govern the subsequent evolution of hybrid genomes. We considered this question in three independent, naturally-occurring hybrid populations formed between swordtail fish species Xiphophorus birchmanni and X. malinche. To this end, we built a fine-scale genetic map and inferred patterns of local ancestry along the genomes of 690 individuals sampled from the three populations. In all three cases, we found hybrid ancestry to be more common in regions of high recombination and where there is linkage to fewer putative targets of selection. These same patterns are also apparent in a reanalysis of human-Neanderthal admixture. Our results lend support to models in which ancestry from the \"minor\" parental species persists only where it is rapidly uncoupled from alleles that are deleterious in hybrids, and show the retention of hybrid ancestry to be at least in part predictable from genomic features. Our analyses further indicate that in swordtail fish, the dominant source of selection on hybrids stems from deleterious combinations of epistatically-interacting alleles.\n\nOne sentence summaryThe persistence of hybrid ancestry is predictable from local recombination rates, in three replicate hybrid populations as well as in humans.

evolutionary biology

Differential Brd4-bound enhancers drive critical sex differences in glioblastoma

Sex can be an important determinant of cancer phenotype, and exploring sex-biased tumor biology holds promise for identifying novel therapeutic targets and new approaches to cancer treatment. In an established isogenic murine model of glioblastoma, we discovered correlated transcriptome-wide sex differences in gene expression, H3K27ac marks, large Brd4-bound enhancer usage, and Brd4 localization to Myc and p53 genomic binding sites. These sex-biased gene expression patterns were also evident in human glioblastoma stem cells (GSCs). These observations led us to hypothesize that Brd4-bound enhancers might underlie sex differences in stem cell function and tumorigenicity in GBM. We found that male and female GBM cells exhibited opposing responses to pharmacological or genetic inhibition of Brd4. Brd4 knockdown or pharmacologic inhibition decreased male GBM cell clonogenicity and in vivo tumorigenesis, while increasing both in female GBM cells. These results were validated in male and female patient-derived GBM cell lines. Furthermore, analysis of the Cancer Therapeutic Response Portal of human GBM samples segregated by sex revealed that male GBM cells are significantly more sensitive to BET inhibitors than are female cells. Thus, for the first time, Brd4 activity is revealed to drive a sex differences in stem cell and tumorigenic phenotype, resulting in diametrically opposite responses to BET inhibition in male and female GBM cells. This has important implications for the clinical evaluation and use of BET inhibitors. SignificanceConsistent sex differences in incidence and outcome have been reported in numerous cancers including brain tumors. GBM, the most common and aggressive primary brain tumor, occurs with higher incidence and shorter survival in males compared to females. Brd4 is essential for regulating transcriptome-wide gene expression and specifying cell identity, including that of GBM. We report that sex-biased Brd4 activity drive sex differences in GBM and render male and female tumor cells differentially sensitive to BET inhibitors. The observed sex differences in BETi treatment strongly indicate that sex differences in disease biology translate into sex differences in therapeutic responses. This has critical implications for clinical use of BET inhibitors further affirming the importance of inclusion of sex as a biological variable.

genomics