Search bioRxivSearch

Biology subjects

Storey, J. D.

Publications and source records attributed to Storey, J. D..

6 recordsLinked to original sources

Real-time in vivo Global Transcriptional Dynamics During Plasmodium falciparum Blood-stage Development

Genome-wide analysis of transcription in the human malaria parasite Plasmodium falciparum has revealed robust variation in steady-state mRNA abundance throughout the 48-hour intraerythrocytic developmental cycle (IDC) suggesting that this process is highly dynamic and tightly regulated. However, the precise timing of mRNA transcription and decay remains poorly understood due to the utilization of methods that only measure total RNA and cannot differentiate between newly transcribed, decaying and stable cellular RNAs. Here we utilize rapid 4-thiouracil (4-TU) incorporation via pyrimidine salvage to specifically label, capture and quantify newly-synthesized P. falciparum RNA transcripts at every hour throughout the IDC following erythrocyte invasion. This high resolution global analysis of the transcriptome captures the timing and rate of transcription for each newly synthesized mRNA in vivo, revealing active transcription throughout all stages of the IDC. To determine the fraction of active transcription and/or transcript stabilization contributing to the total mRNA abundance at each timepoint we have generated a statistical model to fit the data for each gene which reveals varying degrees of transcription and stabilization for each mRNA corresponding to developmental transitions and independent of abundance profile. Finally, our results provide new insight into co-regulation of mRNAs throughout the IDC through regulatory DNA sequence motifs associated with these processes, thereby expanding our understanding of P. falciparum mRNA dynamics.

microbiology

The Functional False Discovery Rate with Applications to Genomics

The false discovery rate measures the proportion of false discoveries among a set of hypothesis tests called significant. This quantity is typically estimated based on p-values or test statistics. In some scenarios, there is additional information available that may be used to more accurately estimate the false discovery rate. We develop a new framework for formulating and estimating false discovery rates and q-values when an additional piece of information, which we call an \"informative variable\", is available. For a given test, the informative variable provides information about the prior probability a null hypothesis is true or the power of that particular test. The false discovery rate is then treated as a function of this informative variable. We consider two applications in genomics. Our first is a genetics of gene expression (eQTL) experiment in yeast where every genetic marker and gene expression trait pair are tested for associations. The informative variable in this case is the distance between each genetic marker and gene. Our second application is to detect differentially expressed genes in an RNA-seq study carried out in mice. The informative variable in this study is the per-gene read depth. The framework we develop is quite general, and it should be useful in a broad range of scientific applications.

genomics

Extending Tests of Hardy-Weinberg Equilibrium to Structured Populations

Testing for Hardy-Weinberg equilibrium (HWE) is an important component in almost all analyses of population genetic data. Genetic markers that violate HWE are often treated as special cases; for example, they may be flagged as possible genotyping errors or they may be investigated more closely for evolutionary signatures of interest. The presence of population structure is one reason why genetic markers may fail a test of HWE. This is problematic because almost all natural populations studied in the modern setting show some degree of structure. Therefore, it is important to be able to detect deviations from HWE for reasons other than structure. To this end, we extend statistical tests of HWE to allow for population structure, which we call a test of \"structural HWE\" (sHWE). Additionally, our new test allows one to automatically choose tuning parameters and identify accurate models of structure. We demonstrate our approach on several important studies, provide theoretical justification for the test, and present empirical evidence for its utility. We anticipate the proposed test will be useful in a broad range of analyses of genome-wide population genetic data.

genomics

A nonparametric estimator of population structure unifying admixture models and principal components analysis

We introduce a simple and computationally efficient method for fitting the admixture model of genetic population structure, called ALStructure. The strategy of ALStructure is to first estimate the low-dimensional linear subspace of the population admixture components and then search for a model within this subspace that is consistent with the admixture models natural probabilistic constraints. Central to this strategy is the observation that all models belonging to this constrained space of solutions are risk-minimizing and have equal likelihood, rendering any additional optimization unnecessary. The low-dimensional linear subspace is estimated through a recently introduced principal components analysis method that is appropriate for genotype data, thereby providing a solution that has both principal components and probabilistic admixture interpretations. Our approach differs fundamentally from other existing methods for estimating admixture, which aim to fit the admixture model directly by searching for parameters that maximize the likelihood function or the posterior probability. We observe that ALStructure typically outperforms existing methods both in accuracy and computational speed under a wide array of simulated and real human genotype datasets. Throughout this work we emphasize that the admixture model is a special case of a much broader class of models for which algorithms similar to ALStructure may be successfully employed.

genomics

FST and kinship for arbitrary population structures I: Generalized definitions

FST is a fundamental measure of genetic differentiation and population structure, currently defined for subdivided populations. FST in practice typically assumes independent, non-overlapping subpopulations, which all split simultaneously from their last common ancestral population so that genetic drift in each subpopulation is probabilistically independent of the other subpopulations. We introduce a generalized FST definition for arbitrary population structures, where individuals may be related in arbitrary ways, allowing for arbitrary probabilistic dependence among individuals. Our definitions are built on identity-by-descent (IBD) probabilities that relate individuals through inbreeding and kinship coefficients. We generalize FST as the mean inbreeding coefficient of the individuals local populations relative to their last common ancestral population. We show that the generalized definition agrees with Wrights original and the independent subpopulation definitions as special cases. We define a novel coancestry model based on \"individual-specific allele frequencies\" and prove that its parameters correspond to probabilistic kinship coefficients. Lastly, we extend the Pritchard-Stephens-Donnelly admixture model in the context of our coancestry model and calculate its FST. To motivate this work, we include a summary of analyses we have carried out in follow-up papers, where our new approach has been applied to simulations and global human data, showcasing the complexity of human population structure, demonstrating our success in estimating kinship and FST, and the shortcomings of existing approaches. The probabilistic framework we introduce here provides a theoretical foundation that extends FST in terms of inbreeding and kinship coefficients to arbitrary population structures, paving the way for new estimators and novel analyses.\n\nNote: This article is Part I of two-part manuscripts. We refer to these in the text as Part I and Part II, respectively.\n\nPart I: Alejandro Ochoa and John D. Storey. \"FST and kinship for arbitrary population structures I: Generalized definitions\". bioRxiv (10.1101/083915) (2019). https://doi.org/10.1101/083915. First published 2016-10-27.\n\nPart II: Alejandro Ochoa and John D. Storey. \"FST and kinship for arbitrary population structures II: Method of moments estimators\". bioRxiv (10.1101/083923) (2019). https://doi.org/10.1101/083923. First published 2016-10-27.

genetics

FST and kinship for arbitrary population structures II: Method of moments estimators

FST and kinship are key parameters often estimated in modern population genetics studies in order to quantitatively characterize structure and relatedness. Kinship matrices have also become a fundamental quantity used in genome-wide association studies and heritability estimation. The most frequently used estimators of FST and kinship are method-of-moments estimators whose accuracies depend strongly on the existence of simple underlying forms of structure, such as the independent subpopulations model of non-overlapping, independently evolving subpopulations. However, modern data sets have revealed that these simple models of structure likely do not hold in many populations, including humans. In this work, we provide new results on the behavior of these estimators in the presence of arbitrarily complex population structures, which results in an improved estimation framework specifically designed for arbitrary population structures. After establishing a framework for assessing bias and consistency of genome-wide estimators, we calculate the accuracy of existing FST and kinship estimators under arbitrary population structures, characterizing biases and estimation challenges unobserved under their originally assumed models of structure. We then present our new approach, which consistently estimates kinship and FST when the minimum kinship value in the dataset is estimated consistently. We illustrate our results using simulated genotypes from an admixture model, constructing a one-dimensional geographic scenario that departs nontrivially from the independent subpopulations model. Our simulations reveal the potential for severe biases in estimates of existing approaches that are overcome by our new framework. This work may significantly improve future analyses that rely on accurate kinship and FST estimates.

genetics