Search bioRxivSearch

Biology subjects

Esteban Burchard

Publications and source records attributed to Esteban Burchard.

4 recordsLinked to original sources

Genome-wide methylation data mirror ancestry information

Genetic data are known to harbor information about human demographics, and genotyping data are commonly used for capturing ancestry information by leveraging genome-wide differences between populations. In contrast, it is not clear to what extent population structure is captured by whole-genome DNA methylation data. We demonstrate, using three large cohort 450K methylation array data sets, that ancestry information signal is mirrored in genome-wide DNA methylation data, and that it can be further isolated more effectively by leveraging the correlation structure of CpGs with cis-located SNPs. Based on these insights, we propose a method, EPISTRUCTURE, for the inference of ancestry from methylation data, without the need for genotype data. EPISTRUCTURE can be used to infer ancestry information of individuals based on their methylation data in the absence of corresponding genetic data. Although genetic data are often collected in epigenetic studies of large cohorts, these are typically not made publicly available, making the application of EPISTRUCTURE especially useful for anyone working on public data. Implementation of EPISTRUCTURE is available in GLINT, our recently released toolset for DNA methylation analysis at: http://glint-epigenetics.readthedocs.io.

Genetics

Dumpster diving in RNA-sequencing to find the source of every last read

High throughput RNA sequencing technologies have provided invaluable research opportunities across distinct scientific domains by producing quantitative readouts of the transcriptional activity of both entire cellular populations and single cells. The majority of RNA-Seq analyses begin by mapping each experimentally produced sequence (i.e., read) to a set of annotated reference sequences for the organism of interest. For both biological and technical reasons, a significant fraction of reads remains unmapped. In this work, we develop Read Origin Protocol (ROP) to discover the source of all reads originating from complex RNA molecules, recombinant T and B cell receptors, and microbial communities. We applied ROP to 8,641 samples across 630 individuals from 54 tissues. A fraction of RNA-Seq data (n=86) was obtained in-house; the remaining data was obtained from the Genotype-Tissue Expression (GTEx v6) project. To generalize the reported number of accounted reads, we also performed ROP analysis on thousands of different, randomly selected, and publicly available RNA-Seq samples in the Sequence Read Archive (SRA). Our approach can account for 99.9% of 1 trillion reads of various read length across the merged dataset (n=10641). Using in-house RNA-Seq data, we show that immune profiles of asthmatic individuals are significantly different from the profiles of control individuals, with decreased average per sample T and B cell receptor diversity. We also show that immune diversity is inversely correlated with microbial load. Our results demonstrate the potential of ROP to exploit unmapped reads in order to better understand the functional mechanisms underlying connections between the immune system, microbiome, human gene expression, and disease etiology. ROP is freely available at https://github.com/smangul1/rop and currently supports human and mouse RNA-Seq reads.

Genomics

Novel Genetic Risk factors for Asthma in African American Children: Precision Medicine and The SAGE II Study.

BackgroundAsthma, an inflammatory disorder of the airways, is the most common chronic disease of children worldwide. There are significant racial/ethnic disparities in asthma prevalence, morbidity and mortality among U.S. children. This trend is mirrored in obesity, which may share genetic and environmental risk factors with asthma. The majority of asthma biomedical research has been performed in populations of European decent.\n\nObjectiveWe sought to identify genetic risk factors for asthma in African American children. We also assessed the generalizability of genetic variants associated with asthma in European and Asian populations to African American children.\n\nMethodsOur study population consisted of 1227 (812 asthma cases, 415 controls) African American children with genome-wide single nucleotide polymorphism (SNP) data. Logistic regression was used to identify associations between SNP genotype and asthma status.\n\nResultsWe identified a novel variant in the PTCHD3 gene that is significantly associated with asthma (rs660498, p = 2.2 x10-7) independent of obesity status. Fewer than 5% of previously reported asthma genetic associations identified in European populations replicated in African Americans.\n\nConclusionsOur identification of novel variants associated with asthma in African American children, coupled with our inability to replicate the majority of findings reported in European Americans, underscores the necessity for including diverse populations in biomedical studies of asthma.

Genetics

An Ancestry Based Approach for Detecting Interactions

IBackgroundEpistasis and gene-environment interactions are known to contribute significantly to variation of complex phenotypes in model organisms. However, their identification in human association studies remains challenging for myriad reasons. In the case of epistatic interactions, the large number of potential interacting sets of genes presents computational, multiple hypothesis correction, and other statistical power issues. In the case of gene-environment interactions, the lack of consistently measured environmental covariates in most disease studies precludes searching for interactions and creates difficulties for replicating studies.\n\nResultsIn this work, we develop a new statistical approach to address these issues that leverages genetic ancestry in admixed populations. We applied our method to gene expression and methylation data from African American and Latino admixed individuals respectively, identifying nine interactions that were significant at p < 5x10-8, we show that two of the interactions in methylation data replicate, and the remaining six are significantly enriched for low p-values (p < 1.8x10-6).\n\nConclusionWe show that genetic ancestry can be a useful proxy for unknown and unmeasured covariates in the search for interaction effects. These results have important implications for our understanding of the genetic architecture of complex traits.

Genetics