Search bioRxivSearch

Biology subjects

de los Campos, G.

Publications and source records attributed to de los Campos, G..

4 recordsLinked to original sources

Quantifying Heterogeneity in the Genetic Architecture of Complex Traits Between Ethnically Diverse Groups using Random Effect Interaction Models

In humans, most genome-wide association studies have been conducted using data from Caucasians and many of the reported findings have not replicated in other populations. This lack of replication may be due to statistical issues (small sample size, confounding) or perhaps more fundamentally to differences in the genetic architecture of traits between ethnically diverse subpopulations. What aspects of the genetic architecture of traits vary between subpopulations and how can this be quantified? We consider studying effect heterogeneity using random-effect Bayesian interaction models. The proposed methodology can be applied using shrinkage and variable selection methods and produces useful information about effect heterogeneity in the form of whole-genome summaries (e.g., SNP-heritability and the average correlation of effects) as well as SNP-specific attributes. Using simulations, we show that the proposed methodology yields (nearly) unbiased estimates of genomic heritability and of the average correlation of effects between groups when the sample size is not too small relative to the number of SNPs used. Subsequently, we used the proposed methodology for the analyses of four complex human traits (standing height, high-density lipoprotein, low-density lipoprotein, and serum urate levels) in European-Americans (EAs) and African-Americans (AAs). The estimated correlations of effects between the two subpopulations was well below unity for all the traits, ranging from 0.73 to 0.50. The extent of effect heterogeneity varied between traits and SNP-sets. Height showed less differences in SNP effects between AAs and EAs whereas HDL, a trait highly influenced by life-style, exhibited greater extent of effect heterogeneity. For all the traits we observed substantial variability in effect heterogeneity across SNPs, suggesting it varies between regions of the genome.

genetics

Imperfect Linkage Disequilibrium Generates Phantom Epistasis (& Perils of Big Data)

The genetic architecture of complex human traits and diseases is affected by large number of possibly interacting genes, but detecting epistatic interactions can be challenging. In the last decade, several studies have alluded to problems that linkage disequilibrium can create when testing for epistatic interactions between DNA markers. However, these problems have not been formalized nor have their consequences been quantified in a precise manner. Here we use a conceptually simple three locus model involving a causal locus and two markers to show that imperfect LD can generate the illusion of epistasis, even when the underlying genetic architecture is purely additive. We describe necessary conditions for such \"phantom epistasis\" to emerge and quantify its relevance using simulations. Our empirical results demonstrate that phantom epistasis can be a very serious problem in GWAS studies (with rejection rates against the additive model greater than 0.2 for nominal p-values of 0.05, even when the model is purely additive). Some studies have sought to avoid this problem by only testing interactions between SNPs with R-sq. <0.1. We show that this threshold is not appropriate and demonstrate that the magnitude of the problem is even greater with large sample size. We conclude that caution must be exercised when interpreting GWAS results derived from very large data sets showing strong evidence in support of epistatic interactions between markers.

genetics

HaploBlocker: Creation of subgroup specific haplotype blocks and libraries

The concept of haplotype blocks has been shown to be useful in genetics. Fields of application range from the detection of regions under positive selection to statistical methods that make use of dimension reduction. We propose a novel approach (\"HaploBlocker\") for defining and inferring haplotype blocks that focuses on linkage instead of the commonly used population-wide measures of linkage disequilibrium (LD) which fail to identify segments shared by individuals in only a subset of the population. We define a haplotype block as a sequence of alleles that has a predefined minimum frequency in the population and only haplotypes with a similar sequence of alleles are considered to be carrying that block, effectively screening a dataset for group-wise identity-by-descent (IBD). Different to most other approaches these blocks are not restricted to shared start or end positions, but can overlap or even contain each other. From these haplotype blocks we construct a haplotype library that represents a large proportion of genetic variability of a population with a limited number of blocks. Our method is implemented in the associated R-package HaploBlocker and provides flexibility to not only optimize the structure of the obtained haplotype library for subsequent analyses (e.g., identification of shared segments between different populations), but is also able to handle datasets of different marker density and genetic diversity. By using haplotype blocks instead of SNPs, local epistatic interactions can be naturally modelled and the reduced number of parameter enables a wide variety of new methods for further genomic analyses. We illustrate our methodology with a dataset comprising 501 doubled haploid lines in a European maize landrace genotyped at 501124 SNPs. With the suggested approach, we identified 2851 haplotype blocks with an average length of 2633 SNPs (compared to 27.8 SNPs per block in HaploView) that together represent 94% of the dataset.\n\nAuthor summaryWhereas it is quite easy to identify segments of shared DNA between pairs of individuals, the problem becomes far more complex when analyzing a population. Especially for livestock and crop populations under strong selection one can observe long and possibly favourable segments that are segregating at high frequency. We propose here an adaptive and flexible approach to identify such segments (\"haplotype blocks\"). The main conceptual difference to other approaches is that we allow haplotype blocks to overlap so that patterns shared by a subset of the population can be mapped adequately. Afterwards, we select a set of those haplotype blocks that form a representation of the whole population (\"haplotype library\"). This haplotype library can be used similar to a SNP-dataset for subsequent genomic approaches with the advantage of a massive reduction of the number of parameters compared to standard haplotyping approaches. Since many breeding goals (e.g. grain yield, milk production) are known to be caused by complex interactions in genomic regions (or even the whole genome) using haplotype blocks instead of single base pairs provides a natural model for local interactions and enables the use of more complex models to incorporate distant interactions between genes, for instance.

genetics

Accurate Genomic Prediction Of Human Height

We construct genomic predictors for heritable and extremely complex human quantitative traits (height, heel bone density, and educational attainment) using modern methods in high dimensional statistics (i.e., machine learning). Replication tests show that these predictors capture, respectively, ~40, 20, and 9 percent of total variance for the three traits. For example, predicted heights correlate ~0.65 with actual height; actual heights of most individuals in validation samples are within a few cm of the prediction. The variance captured for height is comparable to the estimated SNP heritability from GCTA (GREML) analysis, and seems to be close to its asymptotic value (i.e., as sample size goes to infinity), suggesting that we have captured most of the heritability for the SNPs used. Thus, our results resolve the common SNP portion of the \"missing heritability\" problem - i.e., the gap between prediction R-squared and SNP heritability. The ~20k activated SNPs in our height predictor reveal the genetic architecture of human height, at least for common SNPs. Our primary dataset is the UK Biobank cohort, comprised of almost 500k individual genotypes with multiple phenotypes. We also use other datasets and SNPs found in earlier GWAS for out-of-sample validation of our results.

genomics