Search bioRxivSearch

Biology subjects

Xia Shen

Publications and source records attributed to Xia Shen.

7 recordsLinked to original sources

Genomic prediction and estimation of marker interaction effects

The standard models for genomic prediction assume additive polygenic marker effects. For epistatic models including marker interaction effects, the number of effects to be fitted becomes large, which require computational tools tailored specifically for such models. Here, we extend the methods implemented in the R package bigRR so that marker interaction effects can be computed. Simulation results based on marker data from Arabidopsis thaliana show that the inclusion of interaction effects between markers can give a small but significant improvement in genomic predictions. The methods were implemented in the R package EPISbi-gRR available in the bigRR project on R-Forge. The package includes an introductory vignette to the functions available in EPISbigRR.\n\nR package URL: https://r-forge.r-project.org/R/?group_id=1301

Genetics

Genetic regulation of transcriptional variation in natural Arabidopsis thaliana accessions

An increased knowledge of the genetic regulation of expression in Arabidopsis thaliana is likely to provide important insights about the basis of the plants extensive phenotypic variation. Here, we reanalysed two publicly available datasets with genome-wide data on genetic and transcript variation in large collections of natural A. thaliana accessions. Transcripts from more than half of all genes were detected in the leaf of all accessions, and from nearly all annotated genes in at least one accession. Thousands of genes had high transcript levels in some accessions but no transcripts at all in others and this pattern was correlated with the genome-wide genotype. In total, 2,669 eQTL were mapped in the largest population, and 717 of them were replicated in the other population. 646 cis-eQTLs regulated genes that lacked detectable transcripts in some accessions, and for 159 of these we identified one, or several, common structural variants in the populations that were shown to be likely contributors to the lack of detectable RNA-transcripts for these genes. This study thus provides new insights on the overall genetic regulation of global gene-expression diversity in the leaf of natural A. thaliana accessions. Further, it also shows that strong cis-acting polymorphisms, many of which are likely to be structural variations, make important contributions to the transcriptional variation in the worldwide A. thaliana population.

Genetics

Simple multi-trait analysis identifies novel loci associated with growth and obesity measures

The ever-growing genome-wide association studies (GWAS) have revealed widespread pleiotropy. To exploit this, various methods which consider variant association with multiple traits jointly have been developed. However, most effort has been put on improving discovery power: how to replicate and interpret these discovered pleiotropic loci using multivariate methods has yet to be discussed fully. Using only multiple publicly available single-trait GWAS summary statistics, we develop a fast and flexible multi-trait framework that contains modules for (i) multi-trait genetic discovery, (ii) replication of locus pleiotropic profile, and (iii) multi-trait conditional analysis. The procedure is able to handle any level of sample overlap. As an empirical example, we discovered and replicated 23 novel pleiotropic loci for human anthropometry and evaluated their pleiotropic effects on other traits. By applying conditional multivariate analysis on the 23 loci, we discovered and replicated two additional multi-trait associated SNPs. Our results provide empirical evidence that multi-trait analysis allows detection of additional, replicable, highly pleiotropic genetic associations without genotyping additional individuals. The methods are implemented in a free and open source R package MultiABEL.\n\nAuthor summaryBy analyzing large-scale genomic data, geneticists have revealed widespread pleiotropy, i.e. single genetic variation can affect a wide range of complex traits. Methods have been developed to discover such genetic variants. However, we still lack insights into the relevant genetic architecture - What more can we learn from knowing the effects of these genetic variants?\n\nHere, we develop a fast and flexible statistical analysis procedure that includes discovery, replication, and interpretation of pleiotropic effects. The whole analysis pipeline only requires established genetic association study results. We also provide the mathematical theory behind the pleiotropic genetic effects testing.\n\nMost importantly, we show how a replication study can be essential to reveal new biology rather than solely increasing sample size in current genomic studies. For instance, we show that, using our proposed replication strategy, we can detect the difference in genetic effects between studies of different geographical origins.\n\nWe applied the method to the GIANT consortium anthropometric traits to discover new genetic associations, replicated in the UK Biobank, and provided important new insights into growth and obesity.\n\nOur pipeline is implemented in an open-source R package MultiABEL, sufficiently efficient that allows researchers to immediately apply on personal computers in minutes.

Genetics

The "Gini index" in genetics: measuring genetic architecture complexity of quantitative traits

Genetic architecture is a general terminology used and discussed very often in complex traits genetics. It is related to the number of functional loci involved in explaining variation of a complex trait and the distribution of genetic effects across these loci. Understanding the complexity level of the genetic architecture of complex traits is essential for evaluating the potential power of mapping functional loci and prediction of complex traits. However, there has been no quantitative measurement of the genetic architecture complexity, which makes it difficult to link results from genetic data analysis to such terminology. Inspired by the \"Gini index\" for measuring income distribution in economics, I develop a genetic architecture score (\"GA score\") to measure genetic architecture complexity. Simulations indicate that the GA score is an effective measurement of the complexity level of complex traits genetic architecture.

Bioinformatics

Flaw or discovery? Calculating exact p-values for genome-wide association studies in inbred populations

MotivationGenome-wide association studies have been conducted in inbred populations where the sample size is small. The ordinary association p-values and multiple testing correction therefore become questionable, as the detected genetic effect may or may not be due to chance, depending on the minor allele frequency distribution across the genome. Instead of permutation testing, marker-specific false positive rate can be analytically calculated in inbred populations without heterozygotes.\n\nResultsSolutions of exact p-values for genome-wide association studies in inbred populations were derived and implemented. An example is presented to illustrate that the marker-specific experiment-wise p-value varies as the genome-wide minor allele frequency distribution changes. A simulation using real Arabidopsis thaliana genome indicates that the use of exact p-values improves detection power and reduces inflation due to population structure. An analysis of a defense-related case-control phenotype using the exact p-values revealed the causal locus, where markers with higher MAFs had smaller p-values than the top variants with lower MAFs in ordinary genome-wide association analysis.\n\nAvailability and ImplementationProject URL: https://r-forge.r-project.org/projects/statomics/. The R package p.exact: https://r-forge.r-project.org/R/?group_id=2030.\n\nContactxia.shen@ki.se

Bioinformatics

A Genome-Wide Association Analysis Reveals Epistatic Cancellation of Additive Genetic Variance for Root Length in Arabidopsis thaliana

Efforts to identify loci underlying complex traits generally assume that most genetic variance is additive. Here, we examined the genetics of Arabidopsis thaliana root length and found that the narrow-sense heritability for this trait was statistically zero. This low additive genetic variance likely explains why no associations to root length could be found using standard additive-model-based genome-wide association (GWA) approaches. However, the broad-sense heritability for root length was significantly larger, and we therefore also performed an epistatic GWA analysis to map loci contributing to the epistatic genetic variance. This analysis revealed four interacting pairs involving seven chromosomal loci that passed a standard multiple-testing corrected significance threshold. Explorations of the genotype-phenotype maps for these pairs revealed that the detected epistasis cancelled out the additive genetic variance, explaining why these loci were not detected in the additive GWA analysis. Small population sizes, such as in our experiment, increase the risk of identifying false epistatic interactions due to testing for associations with very large numbers of multi-marker genotypes in few phenotyped individuals. Therefore, we estimated the false-positive risk using a new statistical approach that suggested half of the associated pairs to be true positive associations. Our experimental evaluation of candidate genes within the seven associated loci suggests that this estimate is conservative; we identified functional candidate genes that affected root development in four loci that were part of three of the pairs. In summary, statistical epistatic analyses were found to be indispensable for confirming known, and identifying several new, functional candidate genes for root length using a population of wild-collected A. thaliana accessions. We also illustrated how epistatic cancellation of the additive genetic variance resulted in an insignificant narrow-sense, but significant broad-sense heritability that could be dissected into the contributions of several individual loci using a combination of careful statistical epistatic analyses and functional genetic experiments.\n\nAuthor summaryComplex traits, such as many human diseases or climate adaptation and production traits in crops, arise through the action and interaction of many genes and environmental factors. Classic approaches to identify contributing genes generally assume that these factors contribute mainly additive genetic variance. Recent methods, such as genome-wide association studies, often adhere to this additive genetics paradigm. However, additive models of complex traits do not reflect that genes can also contribute with non-additive genetic variance. In this study, we use Arabidopsis thaliana to determine the additive and non-additive genetic contributions to the phenotypic variation in root length. Surprisingly, much of the observed phenotypic variation in root length across genetically divergent strains was explained by epistasis. We mapped seven loci contributing to the epistatic genetic variance and validated four genes in these loci with mutant analysis. For three of these genes, this is their first implication in root development. Together, our results emphasize the importance of considering both non-additive and additive genetic variance when dissecting complex trait variation, in order not to lose sensitivity in genetic analyses.

Genomics

Natural CMT2 variation is associated with genome-wide methylation changes and temperature seasonality

As Arabidopsis thaliana has colonized a wide range of habitats across the world it is an attractive model for studying the genetic mechanisms underlying environmental adaptation. Here, we used public data from two collections of A. thaliana accessions to associate genetic variability at individual loci with differences in climates at the sampling sites. We use a novel method to screen the genome for plastic alleles that tolerate a broader climate range than the major allele. This approach reduces confounding with population structure and increases power compared to standard genome-wide association methods. Sixteen novel loci were found, including an association between Chromomethylase 2 (CMT2) and temperature seasonality where the genome-wide CHH methylation was different for the group of accessions carrying the plastic allele. Cmt2 mutants were shown to be more tolerant to heat-stress, suggesting genetic regulation of epigenetic modifications as a likely mechanism underlying natural adaptation to variable temperatures, potentially through differential allelic plasticity to temperature-stress.\n\nAUTHOR SUMMARYA central problem when studying adaptation to a new environment is the interplay between genetic variation and phenotypic plasticity. Arabidopsis thaliana has colonized a wide range of habitats across the world and it is therefore an attractive model for studying the genetic mechanisms underlying environmental adaptation. Here, we study two collections of A. thaliana accessions from across Eurasia to identify loci associated with differences in climates at the sampling sites. A new genome-wide association analysis method was developed to detect adaptive loci where the alleles tolerate different climate ranges. Sixteen novel such loci were found including a strong association between Chromomethylase 2 (CMT2) and temperature seasonality. The reference allele dominated in areas with less seasonal variability in temperature, and the alternative allele existed in both stable and variable regions. Our results thus link natural variation in CMT2 and epigenetic changes to temperature adaptation. We showed experimentally that plants with a defective CMT2 gene tolerate heat-stress better than plants with a functional gene. Together this strongly suggests a role for genetic regulation of epigenetic modifications in natural adaptation to temperature and illustrates the importance of re-analyses of existing data using new analytical methods to obtain deeper insights into the underlying biology from available data.

Genomics