Search bioRxivSearch

Biology subjects

Stephens, M.

Publications and source records attributed to Stephens, M..

10 recordsLinked to original sources

Discovery and characterization of variance QTLs in human induced pluripotent stem cells

Quantification of gene expression levels at the single cell level has revealed that gene expression can vary substantially even across a population of homogeneous cells. However, it is currently unclear what genomic features control variation in gene expression levels, and whether common genetic variants may impact gene expression variation. Here, we take a genome-wide approach to identify expression variance quantitative trait loci (vQTLs). To this end, we generated single cell RNA-seq (scRNA-seq) data from induced pluripotent stem cells (iPSCs) derived from 53 Yoruba individuals. We collected data for a median of 95 cells per individual and a total of 5,447 single cells, and identified 241 mean expression QTLs (eQTLs) at 10% FDR, of which 82% replicate in bulk RNA-seq data from the same individuals. We further identified 14 vQTLs at 10% FDR, but demonstrate that these can also be explained as effects on mean expression. Our study suggests that dispersion QTLs (dQTLs) which could alter the variance of expression independently of the mean can have larger fold changes, but explain less phenotypic variance than eQTLs. We estimate 424 individuals as a lower bound to achieve 80% power to detect the strongest dQTLs in iPSCs. These results will guide the design of future studies on understanding the genetic control of gene expression variance.\n\nAuthor summaryCommon genetic variation can alter the level of average gene expression in human tissues, and through changes in gene expression have downstream consequences on cell function, human development, and human disease. However, human tissues are composed of many cells, each with its own level of gene expression. With advances in single cell sequencing technologies, we can now go beyond simply measuring the average level of gene expression in a tissue sample and directly measure cell-to-cell variance in gene expression. We hypothesized that genetic variation could also alter gene expression variance, potentially revealing new insights into human development and disease. To test this hypothesis, we used single cell RNA sequencing to directly measure gene expression variance in multiple individuals, and then associated the gene expression variance with genetic variation in those same individuals. Our results suggest that effects on gene expression variance are smaller than effects on mean expression, relative to how much the phenotypes vary between individuals, and will require much larger studies than previously thought to detect.

genomics

CorShrink : Empirical Bayes shrinkage estimation of correlations, with applications

Estimation of correlation matrices and correlations among variables is a ubiquitous problem in statistics. In many cases - especially when the number of observations is small relative to the number of variables - some kind of shrinkage or regularization is necessary to improve estimation accuracy. Here, we propose an Empirical Bayes shrinkage approach, CorShrink, which adaptively learns how much to shrink correlations by combining information across all pairs of variables. One key feature of CorShrink, which distinguishes it from most existing methods, is its flexibility in dealing with missing data. Indeed, CorShrink explicitly accounts for varying amounts of missingness among pairs of variables. Numerical studies suggest CorShrink is competitive with other popular correlation shrinkage methods, even when there is no missing data. We illustrate CorShrink on gene expression data from GTEx project, which suffers from extensive missing observations, and where existing methods struggle. We also illustrate its flexibility by applying it to estimate cosine similarities between word vectors from word2vec models, thereby generating more accurate word similarity rankings.

genetics

Model-based analysis of positive selection significantly expands the list of cancer driver genes, including RNA methyltransferases

Identifying driver genes is a central problem in cancer biology, and many methods have been developed to identify driver genes from somatic mutation data. However, existing methods either lack explicit statistical models, or rely on very simple models that do not capture complex features in somatic mutations of driver genes. Here, we present driverMAPS (Model-based Analysis of Positive Selection), a more comprehensive model-based approach to driver gene identification. This new method explicitly models, at the single-base level, the effects of positive selection in cancer driver genes as well as highly heterogeneous background mutational process. Its selection model captures elevated mutation rates in functionally important sites using multiple external annotations, as well as spatial clustering of mutations. Its background mutation model accounts for both known covariates and unexplained local variation. Simulations under realistic evolutionary models demonstrate that driverMAPS greatly improves the power of driver gene detection over state-of-the-art approaches. Applying driverMAPS to TCGA data across 20 tumor types identified 159 new potential driver genes. Cross-referencing this list with data from external sources strongly supports these findings. The novel genes include the mRNA methytransferases METTL3-METTL14, and we experimentally validated METTL3 as a potential tumor suppressor gene in bladder cancer. Our results thus provide strong support to the emerging hypothesis that mRNA modification is an important biological process underlying tumorigenesis.

genomics

Estimating recent migration and population size surfaces

In many species a fundamental feature of genetic diversity is that genetic similarity decays with geographic distance; however, this relationship is often complex, and may vary across space and time. Methods to uncover and visualize such relationships have widespread use for analyses in molecular ecology, conservation genetics, evolutionary genetics, and human genetics. While several frameworks exist, a promising approach is to infer maps of how migration rates vary across geographic space. Such maps could, in principle, be estimated across time to reveal the full complexity of population histories. Here, we take a step in this direction: we present a method to infer separate maps of population sizes and migration rates for different time periods from a matrix of genetic similarity between every pair of individuals. Specifically, genetic similarity is measured by counting the number of long segments of haplotype sharing (also known as identity-by-descent tracts). By varying the length of these segments we obtain parameter estimates for qualitatively different time periods. Using simulations, we show that the method can reveal time-varying migration rates and population sizes, including changes that are not detectable when ignoring haplotypic structure. We apply the method to a dataset of contemporary European individuals (POPRES), and provide an integrated analysis of recent population structure and growth over the last ~3,000 years in Europe. Software implementing the methods is available at https://github.com/halasadi/MAPS.

bioinformatics

Inference and visualization of DNA damage patterns using a grade of membership model

Quality control plays a major role in the analysis of ancient DNA (aDNA). One key step in this quality control is assessment of DNA damage: aDNA contains unique signatures of DNA damage that distinguish it from modern DNA, and so analyses of damage patterns can help confirm that DNA sequences obtained are from endogenous aDNA rather than from modern contamination. Predominant signatures of DNA damage include a high frequency of cytosine to thymine substitutions (C-to-T) at the ends of fragments, and elevated rates of purines (A & G) before the 5 strand-breaks. Existing QC procedures help assess damage by simply plotting for each sample, the C-to-T mismatch rate along the read and the composition of bases before the 5 strand-breaks. Here we present a more flexible and comprehensive model-based approach to infer and visualize damage patterns in aDNA, implemented in an R package aRchaic. This approach is based on a \"grade of membership\" model (also known as \"admixture\" or \"topic\" model) in which each sample has an estimated grade of membership in each of K damage profiles that are estimated from the data. We illustrate aRchaic on data from several aDNA studies and modern individuals from 1000 Genomes Project Consortium (2012). Here, aRchaic clearly distinguishes modern from ancient samples irrespective of DNA extraction, lab and sequencing protocols. Additionally, through an in-silico contamination experiment, we show that the aRchaic grades of membership reflect relative levels of exogenous modern contamination. Together, the outputs of aRchaic provide a concise visual summary of DNA damage patterns, as well as other processes generating mismatches in the data. Availability: aRchaic is available for download from https://www.github.com/kkdey/aRchaic.\n\nContact: halasadi@uchicago.edu, kkdey@uchicago.edu

bioinformatics

Harnessing Empirical Bayes and Mendelian Segregation for Genotyping Autopolyploids from Messy Sequencing Data

Detecting and quantifying the differences in individual genomes (i.e. genotyping), plays a fundamental role in most modern bioinformatics pipelines. Many scientists now use reduced representation next-generation sequencing (NGS) approaches for genotyping. Genotyping diploid individuals using NGS is a well-studied field and similar methods for polyploid individuals are just emerging. However, there are many aspects of NGS data, particularly in polyploids, that remain unexplored by most methods. We provide two main contributions in this paper: (i) We draw attention to, and then model, common aspects of NGS data: sequencing error, allelic bias, overdispersion, and outlying observations. (ii) Many datasets feature related individuals, and so we use the structure of Mendelian segregation to build an empirical Bayes approach for genotyping polyploid individuals. We assess the accuracy of our method in simulations and apply it to a dataset of hexaploid sweet potatoes (Ipomoea batatas). An R package implementing our method is available at https://github.com/dcgerard/updog.

bioinformatics

A new sequence logo plot to highlight enrichment and depletion

BackgroundSequence logo plots have become a standard graphical tool for visualizing sequence motifs in DNA, RNA or protein sequences. However standard logo plots primarily highlight enrichment of symbols, and may fail to highlight interesting depletions. Current alternatives that try to highlight depletion often produce visually cluttered logos.\n\nResultsWe introduce a new sequence logo plot, the EDLogo plot, that highlights both enrichment and depletion, while minimizing visual clutter. We provide an easy-to-use and highly customizable R package Logolas to produce a range of logo plots, including EDLogo plots. This software also allows elements in the logo plot to be strings of characters, rather than a single character, extending the range of applications beyond the usual DNA, RNA or protein sequences. We illustrate our methods and software on applications to transcription factor binding site motifs, protein sequence alignments and cancer mutation signature profiles.\n\nConclusionOur new EDLogo plots, and flexible software implementation, can help data analysts visualize both enrichment and depletion of characters (DNA sequence bases, amino acids, etc) across a wide range of applications.

genetics

A large-scale genome-wide enrichment analysis identifies new trait-associated genes, pathways and tissues across 31 human phenotypes

Genome-wide association studies (GWAS) aim to identify genetic factors that are associated with complex traits. Standard analyses test individual genetic variants, one at a time, for association with a trait. However, variant-level associations are hard to identify (because of small effects) and can be difficult to interpret biologically. \"Enrichment analyses\" help address both these problems by focusing on sets of biologically-related variants. Here we introduce a new model-based enrichment analysis method that requires only GWAS summary statistics, and has several advantages over existing methods. Applying this method to interrogate 3,913 biological pathways and 113 tissue-based gene sets in 31 human phenotypes identifies many previously-unreported enrichments. These include enrichments of the endochondral ossification pathway for adult height, the NFAT-dependent transcription pathway for rheumatoid arthritis, brain-related genes for coronary artery disease, and liver-related genes for late-onset Alzheimers disease. A key feature of our method is that inferred enrichments automatically help identify new trait-associated genes. For example, accounting for enrichment in lipid transport genes yields strong evidence for association between MTTP and low-density lipoprotein levels, whereas conventional analyses of the same data found no significant variants near this gene.

genomics

Silencing Of Transposable Elements May Not Be A Major Driver Of Regulatory Evolution In Primate Induced Pluripotent Stem Cells

Transposable elements (TEs) comprise a substantial proportion of primate genomes. The regulatory potential of TEs can result in deleterious effects, especially during development. It has been suggested that, in pluripotent stem cells, TEs are targeted for silencing by KRAB-ZNF proteins, which recruit the TRIM28-SETDB1 complex, to deposit the repressive histone modification H3K9me3. TEs, in turn, can acquire mutations that allow them to evade detection by the host, and hence KRAB-ZNF proteins need to rapidly evolve to counteract them. To investigate the short-term evolution of TE silencing, we profiled the genome-wide distribution of H3K9me3 in induced pluripotent stem cells from ten human and seven chimpanzee individuals. We performed chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) for H3K9me3, as well as total RNA sequencing. We focused specifically on cross-species H3K9me3 ChIP-seq data that mapped to four million orthologous TEs. We found that, depending on the TE class, 10-60% of elements are marked by H3K9me3, with SVA, LTR and LINE elements marked most frequently. We found little evidence of inter-species differences in TE silencing, with as many as 80% of orthologous, putatively silenced, TEs marked at similar levels in humans and chimpanzees. Our data suggest limited species-specificity of TE silencing across six million years of primate evolution. Interestingly, the minority of TEs enriched for H3K9me3 in one species are not more likely to be associated with gene expression divergence of nearby orthologous genes. We conclude that orthologous TEs may not play a major role in driving gene regulatory divergence between humans and chimpanzees.

genomics

Flexible statistical methods for estimating and testing effects in genomic studies with multiple conditions

We introduce new statistical methods for analyzing genomic datasets that measure many effects in many conditions (e.g., gene expression changes under many treatments). These new methods improve on existing methods by allowing for arbitrary correlations in effect sizes among conditions. This flexible approach increases power, improves effect estimates, and allows for more quantitative assessments of effect-size heterogeneity compared to simple "shared/condition-specific" assessments. We illustrate these features through an analysis of locally-acting variants associated with gene expression ("cis eQTLs") in 44 human tissues. Our analysis identifies more eQTLs than existing approaches, consistent with improved power. We show that while genetic effects on expression are extensively shared among tissues, effect sizes can still vary greatly among tissues. Some shared eQTLs show stronger effects in subsets of biologically related tissues (e.g., brain-related tissues), or in only one tissue (e.g., testis). Our methods are widely applicable, computationally tractable for many conditions, and available online.

genomics