Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

An ultra-dense haploid genetic map for evaluating the highly fragmented genome assembly of Norway spruce (Picea abies)

Norway spruce (Picea abies (L.) Karst.) is a conifer species of substanital economic and ecological importance. In common with most conifers, the P. abies genome is very large ([~]20 Gbp) and contains a high fraction of repetitive DNA. The current P. abies genome assembly (v1.0) covers approximately 60% of the total genome size but is highly fragmented, consisting of >10 million scaffolds. The genome annotation contains 66,632 gene models that are at least partially validated (www.congenie.org), however, the fragmented nature of the assembly means that there is currently little information available on how these genes are physically distributed over the 12 P. abies chromosomes. By creating an ultra-dense genetic linkage map, we anchored and ordered scaffolds into linkage groups, which complements the fine-scale information available in assembly contigs. Our ultra-dense haploid consensus genetic map consists of 21,056 markers derived from 14,336 scaffolds that contain 17,079 gene models (25.6% of the validated gene models) that we have anchored to the 12 linkage groups. We used data from three independent component maps, as well as comparisons with previously published Picea maps to evaluate the accuracy and marker ordering of the linkage groups. We demonstrate that approximately 3.8% of the anchored scaffolds and 1.6% of the gene models covered by the consensus map have likely assembly errors as they contain genetic markers that map to different regions within or between linkage groups. We further evaluate the utility of the genetic map for the conifer research community by using an independent data set of unrelated individuals to assess genome-wide variation in genetic diversity using the genomic regions anchored to linkage groups. The results show that our map is sufficiently dense to enable detailed evolutionary analyses across the P. abies genome.

genetics

Brain scans from 21297 individuals reveal the genetic architecture of hippocampal subfield volumes

The hippocampus is a heterogeneous structure, comprising histologically distinguishable subfields. These subfields are differentially involved in memory consolidation, spatial navigation and pattern separation, complex functions often impaired in individuals with brain disorders characterized by reduced hippocampal volume, including Alzheimers disease (AD) and schizophrenia. Given the structural and functional heterogeneity of the hippocampal formation, we sought to characterize the subfields genetic architecture. T1-weighted brain scans (n=21297, 16 cohorts) were processed with the hippocampal subfields algorithm in FreeSurfer v6.0. We ran a genome-wide association analysis on each subfield, covarying for total hippocampal volume. We further calculated the single nucleotide polymorphism (SNP)-based heritability of twelve subfields, as well as their genetic correlation with each other, with other structural brain features, and with AD and schizophrenia. All outcome measures were corrected for age, sex, and intracranial volume. We found 15 unique genome-wide significant loci across six subfields, of which eight had not been previously linked to the hippocampus. Top SNPs were mapped to genes associated with neuronal differentiation, locomotor behaviour, schizophrenia and AD. The volumes of all the subfields were estimated to be heritable (h2 from .14 to .27, all p< 1x10-16) and clustered together based on their genetic correlations compared to other structural brain features. There was also evidence of genetic overlap of subicular subfield volumes with schizophrenia. We conclude that hippocampal subfields have partly distinct genetic determinants associated with specific biological processes and traits. Taking into account this specificity may increase our understanding of hippocampal neurobiology and associated pathologies.

genetics

Inferring the ancestry of parents and grandparents from genetic data

Inference of admixture proportions is a classical statistical problem in population genetics. Standard methods implicitly assume that both parents of an individual have the same admixture fraction. However, this is rarely the case in real data. In this paper we show that the distribution of admixture tract lengths in a genome contains information about the admixture proportions of the ancestors of an individual. We develop a Hidden Markov Model (HMM) framework for estimating the admixture proportions of the immediate ancestors of an individual, i.e. a type of decomposition of an individuals admixture proportions into further subsets of ancestral proportions in the ancestors. Based on a genealogical model for admixture tracts, we develop an efficient algorithm for computing the sampling probability of the genome from a single individual, as a function of the admixture proportions of the ancestors of this individual. This allows us to perform probabilistic inference of admixture proportions of ancestors only using the genome of an extant individual. We perform extensive simulations to quantify the error in the estimation of ancestral admixture proportions under various conditions. To illustrate the utility of the method, we apply it to real genetic data. Author summaryAncestry inference is an important problem in genetics and is used commercially by a number of companies affecting millions of consumers of genetic ancestry tests. In this paper, we show that it is possible, not only to estimate the ancestry fractions of an individual, but also, with some uncertainty, to estimate the ancestry fractions of an individuals ancestors. For example, if an individual traces his/her ancestry 50% to Asia and 50% to Europe, it is possible to distinguish between the individual having two parents that each are 50:50 composites of Asian and European ancestry, or one parent from Asia and one from Europe. It is likewise also possible to make inferences about grandparents. We present a computationally efficient method for making such inferences called PedMix. PedMix is based on a probabilistic model for the descendant and the recent ancestors. PedMix infers admixture proportions of recent ancestors (parents, grandparents or even great grandparents) using whole-genome genetic variation data from a focal individual. Results on both simulated and real data show that PedMix performs reasonably well in most scenarios.

genetics

Common genetic variants contribute to risk of rare severe neurodevelopmental disorders

There are thousands of rare human disorders caused by a single deleterious, protein-coding genetic variant 1. However, patients with the same genetic defect can have different clinical presentation 2-4, and some individuals carrying known disease-causing variants can appear unaffected 5. What explains these differences? Here, we show in a cohort of 6,987 children with heterogeneous severe neurodevelopmental disorders expected to be almost entirely monogenic that 7.7% of variance in risk is attributable to inherited common genetic variation. We replicated this genome wide common variant burden by showing that it is over-transmitted from parents to children in an independent sample of 728 trios from the same cohort. Our common variant signal is significantly positively correlated with genetic predisposition to fewer years of schooling, decreased intelligence, and risk of schizophrenia. We found that common variant risk was not significantly different between individuals with and without a known protein-coding diagnostic variant, suggesting that common variant risk is not confined to patients without a monogenic diagnosis. In addition, previously published common variant scores for autism, height, birth weight, and intracranial volume were all correlated with those traits within our cohort, suggesting that phenotypic expression in individuals with monogenic disorders is affected by the same variants as the general population. Our results demonstrate that common genetic variation affects both overall risk and clinical presentation in disorders typically considered to be monogenic.

genetics

Bayesian Multi-SNP Genetic Association Analysis: Control of FDR and Use of Summary Statistics

Multi-SNP genetic association analysis has become increasingly important in analyzing data from genome-wide association studies (GWASs) and molecular quantitative trait loci (QTL) mapping studies. In this paper, we propose novel computational approaches to address two outstanding issues in Bayesian multi-SNP genetic association analysis: namely, the control of false positive discoveries of identified association signals and the maximization of the efficiency of statistical inference by utilizing summary statistics. Quantifying the strength and uncertainty of genetic association signals has been a long-standing theme in statistical genetics. However, there is a lack of formal statistical procedures that can rigorously control type I errors in multi-SNP analysis. We propose an intuitive hierarchical representation of genetic association signals based on Bayesian posterior probabilities, which subsequently enables rigorous control of false discovery rate (FDR) and construction of Bayesian credible sets. From the perspective of statistical data reduction, we examine the computational approaches of multi-SNP analysis using z-statistics from single-SNP association testing and conclude that they likely yield conservative results comparing to using individual-level data. Built on this result, we propose a set of sufficient summary statistics that can lead to identical results as individual-level data without sacrificing power. Our novel computational approaches are implemented in the software package, DAP-G (https://github.com/xqwen/dap), which applies to both GWASs and genome-wide molecular QTL mapping studies. It is highly computationally efficient and approximately 20 times faster than the state-of-the-art implementation of Bayesian multi-SNP analysis software. We demonstrate the proposed computational approaches using carefully constructed simulation studies and illustrate a complete workflow for multi-SNP analysis of cis expression quantitative trait loci using the whole blood data from the GTEx project.

genetics

Trans-ethnic polygenic analysis supports genetic overlaps of lumbar disc degeneration with height, body mass index, and bone mineral density

Lumbar disc degeneration (LDD) is age-related break-down in the fibrocartilaginous joints between lumbar vertebrae. It is a major cause of low back pain and is conventionally assessed by magnetic resonance imaging (MRI). Like most other complex traits, LDD is likely polygenic and influenced by both genetic and environmental factors. However, genome-wide association studies (GWASs) of LDD have uncovered few susceptibility loci due to the limited sample size. Previous epidemiology studies of LDD also reported multiple heritable risk factors, including height, body mass index (BMI), bone mineral density (BMD), lipid levels, etc. Genetics can help elucidate causality between traits and suggest loci with pleiotropic effects. One such approach is polygenic score (PGS) which summarizes the effect of multiple variants by the summation of alleles weighted by estimated effects from GWAS. To investigate genetic overlaps of LDD and related heritable risk factors, we calculated the PGS of height, BMI, BMD and lipid levels in a Chinese population-based cohort with spine MRI examination and a Japanese case-control cohort of lumbar disc herniation (LDH) requiring surgery. Because most large-scale GWASs were done in European populations, PGS of corresponding traits were created using weights from European GWASs. We calibrated their prediction performance in independent Chinese samples, then tested associations with MRI-derived LDD scores and LDH affection status. The PGS of height, BMI, BMD and lipid levels were strongly associated with respective phenotypes in Chinese, although phenotype variances explained were lower than in Europeans which would reduce the power to detect genetic overlaps. Despite of this, the PGS of BMI and lumbar spine BMD were significantly associated with LDD scores; and the PGS of height was associated with the increased the liability of LDH. Furthermore, linkage disequilibrium score regression suggested that, osteoarthritis, another degenerative disorder that shares common features with LDD, also showed genetic correlations with height, BMI and BMD. The findings suggest a common key contribution of biomechanical stress to the pathogenesis of LDD and will direct the future search for pleiotropic genes.

genetics

A novel method for systematic genetic analysis and visualization of phenotypic heterogeneity applied to orofacial clefts

Phenotypic heterogeneity is a hallmark of complex traits, and genetic studies may focus on the trait as a whole or on individual subgroups. For example, in orofacial clefting (OFC), three subtypes - cleft lip (CL), cleft lip and palate (CLP), and cleft palate (CP) have been studied separately and in combination. It is more challenging, however, to dissect the genetic architecture and describe how a given locus may be contributing to distinct subtypes of a trait. We developed a framework for quantifying and interpreting evidence of subtype-specific or shared genetic effects in complex traits. We applied this technique to create a \"cleft map\" of the association of 30 genetic loci with three OFC subtypes. In addition to new associations, we found loci with subtype-specific effects (e.g., GRHL3 (CP), WNT5A (CLP)), as well as loci associated with two or all three subtypes. We cross-referenced these results with mouse craniofacial gene expression datasets, which identified promising candidate genes. However, we found no strong correlation between OFC subtypes and expression patterns. In aggregate, the cleft map revealed neither subtype-specific nor shared genetic effects operate in isolation in OFC architecture. Our approach can be easily applied to any complex trait with distinct phenotypic subgroups.

genetics

Using Topic Modeling via Non-negative Matrix Factorization to Identify Relationships between Genetic Variants and Disease Phenotypes: A Case Study of Lipoprotein(a) (LPA)

Genome-wide and phenome-wide association studies are commonly used to identify important relationships between genetic variants and phenotypes. Most of these studies have treated diseases as independent variables and suffered from heavy multiple adjustment burdens due to the large number of genetic variants and disease phenotypes. In this study, we propose using topic modeling via non-negative matrix factorization (NMF) for identifying associations between disease phenotypes and genetic variants. Topic modeling is an unsupervised machine learning approach that can be used to learn the semantic patterns from electronic health record data. We chose rs10455872 in LPA as the predictor since it has been shown to be associated with increased risk of hyperlipidemia and cardiovascular diseases (CVD). Using data of 12,759 individuals from the biobank at Vanderbilt University Medical Center, we trained a topic model using NMF from 1,853 distinct phecodes extracted from the cohorts electronic health records and generated six topics. We quantified their associations with rs10455872 in LPA. Topics indicating CVD had positive correlations with rs10455872 (P < 0.001), replicating a previous finding. We also identified a negative correlation between LPA and a topic representing lung cancer (P < 0.001). Our results demonstrate the applicability of topic modeling in exploring the relationship between the genome and clinical diseases.\n\nAuthor summaryIdentifying the clinical associations of genetic variants remains crucial in understanding how the human genome modulates disease risk. Traditional phenome-wide association studies consider each disease phenotype as an independent variable, however, diseases often present as complex clusters of comorbid conditions. In this study, we propose using topic modeling to model electronic health record data as a mixture of topics (e.g., disease clusters or relevant comorbidities) and testing associations between topics and genetic variants. Our results demonstrated the feasibility of using topic modeling to replicate and discover novel associations between the human genome and clinical diseases.

genetics

Genome-wide association study reveals sex-specific genetic architecture of facial attractiveness

Facial attractiveness is a complex human trait of great interest in both academia and industry. Literature on sociological and phenotypic factors associated with facial attractiveness is rich, but its genetic basis is poorly understood. In this paper, we conducted a genome-wide association study to discover genetic variants associated with facial attractiveness using 3,928 samples in the Wisconsin Longitudinal Study. We identified two genome-wide significant loci and highlighted a handful of candidate genes, many of which are specifically expressed in human tissues involved in reproduction and hormone synthesis. Additionally, facial attractiveness showed strong and negative genetic correlations with BMI in females and with blood lipids in males. Our analysis also suggested sex-specific selection pressure on variants associated with lower male attractiveness. These results revealed sex-specific genetic architecture of facial attractiveness and provided fundamental new insights into its genetic basis.

genetics

Expanded genetic landscape of chronic obstructive pulmonary disease reveals heterogeneous cell type and phenotype associations

Chronic obstructive pulmonary disease (COPD) is the leading cause of respiratory mortality worldwide. Genetic risk loci provide novel insights into disease pathogenesis. To broaden COPD genetic risk loci discovery and identify cell type and phenotype associations we performed a genome-wide association study in 35,735 cases and 222,076 controls from the UK Biobank and additional studies from the International COPD Genetics Consortium. We identified 82 loci with P value < 5x10-8; 47 were previously described in association with either COPD or population-based lung function. Of the remaining 35 novel loci, 13 were associated with lung function in 79,055 individuals from the SpiroMeta consortium. Using gene expression and regulation data, we identified enrichment for loci in lung tissue, smooth muscle and alveolar type II cells. We found 9 shared genomic regions between COPD and asthma and 5 between COPD and pulmonary fibrosis. COPD genetic risk loci clustered into groups of quantitative imaging features and comorbidity associations. Our analyses provide further support to the genetic susceptibility and heterogeneity of COPD.

genetics

Genetics & the Geography of Health, Behavior, and Attainment

Peoples life chances can be predicted by their neighborhoods. This observation is driving efforts to improve lives by changing neighborhoods. Some neighborhood effects may be causal, supporting neighborhood-level interventions. Other neighborhood effects may reflect selection of families with different characteristics into different neighborhoods, supporting interventions that target families/individuals directly. To test how selection affects different neighborhood-linked problems, we linked neighborhood data with genetic, health, and social-outcome data for >7,000 European-descent UK and US young people in the E-Risk and Add Health Studies. We tested selection/concentration of genetic risks for obesity, schizophrenia, teen-pregnancy, and poor educational outcomes in high-risk neighborhoods, including genetic analysis of neighborhood mobility. Findings argue against genetic selection/concentration as an explanation for neighborhood gradients in obesity and mental-health problems, suggesting neighborhoods may be causal. In contrast, modest genetic selection/concentration was evident for teen-pregnancy and poor educational outcomes, suggesting neighborhood effects for these outcomes should be interpreted with care.

genetics

Low genetic variation is associated with low mutation rate in the giant duckweed

Mutation rate and effective population size (Ne) jointly determine intraspecific genetic diversity, but the role of mutation rate is often ignored. We investigate genetic diversity, spontaneous mutation rate and Ne in the giant duckweed (Spirodela polyrhiza). Despite its large census population size, whole-genome sequencing of 68 globally sampled individuals revealed extremely low within-species genetic diversity. Assessed under natural conditions, the genome-wide spontaneous mutation rate is at least seven times lower than estimates made for other multicellular eukaryotes, whereas Ne is large. These results demonstrate that low genetic diversity can be associated with large-Ne species, where selection can reduce mutation rates to very low levels, and accurate estimates of mutation rate can help to explain seemingly counterintuitive patterns of genome-wide variation.\n\nOne Sentence SummaryThe low-down on a tiny plant: extremely low genetic diversity in an aquatic plant is associated with its exceptionally low mutation rate.

genetics

The genetic architecture of the human cerebral cortex

The cerebral cortex underlies our complex cognitive capabilities, yet we know little about the specific genetic loci influencing human cortical structure. To identify genetic variants, including structural variants, impacting cortical structure, we conducted a genome-wide association meta-analysis of brain MRI data from 51,662 individuals. We analysed the surface area and average thickness of the whole cortex and 34 regions with known functional specialisations. We identified 255 nominally significant loci (P [&le;] 5 x 10-8); 199 survived multiple testing correction (P [&le;] 8.3 x 10-10; 187 surface area; 12 thickness). We found significant enrichment for loci influencing total surface area within regulatory elements active during prenatal cortical development, supporting the radial unit hypothesis. Loci impacting regional surface area cluster near genes in Wnt signalling pathways, known to influence progenitor expansion and areal identity. Variation in cortical structure is genetically correlated with cognitive function, Parkinsons disease, insomnia, depression and ADHD.\n\nOne Sentence SummaryCommon genetic variation is associated with inter-individual variation in the structure of the human cortex, both globally and within specific regions, and is shared with genetic risk factors for some neuropsychiatric disorders.

genetics

Parkinson disease age of onset GWAS: defining heritability, genetic loci and a-synuclein mechanisms

Increasing evidence supports an extensive and complex genetic contribution to Parkinsons disease (PD). Previous genome-wide association studies (GWAS) have shed light on the genetic basis of risk for this disease. However, the genetic determinants of PD age of onset are largely unknown. Here we performed an age of onset GWAS based on 28,568 PD cases. We estimated that the heritability of PD age of onset due to common genetic variation was ~0.11, lower than the overall heritability of risk for PD (~0.27) likely in part because of the subjective nature of this measure. We found two genome-wide significant association signals, one at SNCA and the other a protein-coding variant in TMEM175, both of which are known PD risk loci and a Bonferroni corrected significant effect at other known PD risk loci, INPP5F/BAG3, FAM47E/SCARB2, and MCCC1. In addition, we identified that GBA coding variant carriers had an earlier age of onset compared to non-carriers. Notably, SNCA, TMEM175, SCARB2, BAG3 and GBA have all been shown to either directly influence alpha-synuclein aggregation or are implicated in alpha-synuclein aggregation pathways. Remarkably, other well-established PD risk loci such as GCH1, MAPT and RAB7L1/NUCKS1 (PARK16) did not show a significant effect on age of onset of PD. While for some loci, this may be a measure of power, this is clearly not the case for the MAPT locus; thus genetic variability at this locus influences whether but not when an individual develops disease. We believe this is an important mechanistic and therapeutic distinction. Furthermore, these data support a model in which alpha-synuclein and lysosomal mechanisms impact not only PD risk but also age of disease onset and highlights that therapies that target alpha-synuclein aggregation are more likely to be disease-modifying than therapies targeting other pathways.

genetics

The genetic relationship between female reproductive traits and six psychiatric disorders

Female reproductive behaviors have an important implication in evolutionary fitness and health of offspring. Previous studies have shown that age at first birth of women (AFB) is genetically associated with schizophrenia (SCZ). However, for most other psychiatric disorders and reproductive traits, the latent shared genetic architecture is largely unknown. Here we used the second wave of UK Biobank data (N=220,685) to evaluate the association between five female reproductive traits and polygenetic risk scores (PRS) projected from genome-wide association study summary statistics of six psychiatric disorders (N=429,178). We found that the PRS of attention-deficit/hyperactivity disorder (ADHD) were strongly associated with AFB (genetic correlation of -0.68 {+/-} 0.03 with p-value = 1.86E-89), age at first sexual intercourse (AFS) (-0.56 {+/-} 0.03 with p-value = 3.42E-60), number of live births (NLB) (0.36 {+/-} 0.04 with p-value = 4.01E-17) and age at menopause (-0.27 {+/-} 0.04 with p-value = 5.71E-13). There were also robustly significant associations between the PRS of eating disorder (ED) and AFB (genetic correlation of 0.35 {+/-} 0.06), ED and AFS (0.19 0.06), Major depressive disorder (MDD) and AFB (-0.27 {+/-} 0.07), MDD and AFS (- 0.27 {+/-} 0.03) and SCZ and AFS (-0.10 {+/-} 0.03). Our findings reveal the shared genetic architecture between the five reproductive traits in women and six psychiatric disorders, which have a potential implication that helps to improve reproductive health in women, hence better child outcomes. Our findings may also explain, at least in part, an evolutionary hypothesis that causal mutations underlying psychiatric disorders have positive effects on reproductive success.

genetics

Polygenic prediction of breast cancer: comparison of genetic predictors and implications for screening

BackgroundPublished genetic risk scores for breast cancer (BC) so far have been based on a relatively small number of markers and are not necessarily using the full potential of large-scale Genome-Wide Association Studies. This study aims to identify an efficient polygenic predictor for BC based on best available evidence and to assess its potential for personalized risk prediction and screening strategies.\n\nMethodsFour different genetic risk scores (two already published and two newly developed) and their combinations (metaGRS) are compared in the subsets of two population-based biobank cohorts: the UK Biobank (UKBB, 3157 BC cases, 43,827 controls) and Estonian Biobank (EstBB, 317 prevalent and 308 incident BC cases in 32,557 women). In addition, correlations between different genetic risk scores and their associations with BC risk factors are studied in both cohorts.\n\nResultsThe metaGRS that combines two genetic risk scores (metaGRS2 - based on 75 and 898 Single Nucleotide Polymorphisms, respectively) has the strongest association with prevalent BC status in both cohorts. One standard deviation difference in the metaGRS2 corresponds to an Odds Ratio = 1.6 (95% CI 1.54 to 1.66, p = 9.7*10-135) in the UK Biobank and accounting for family history marginally attenuates the effect (Odds Ratio = 1.58, 95% CI 1.53 to 1.64, p = 9.1*10-129). In the EstBB cohort, the hazard ratio of incident BC for the women in the top 5% of the metaGRS2 compared to women in the lowest 50% is 4.2 (95% CI 2.8 to 6.2, p = 8.1*10-13). The different GRSs are only moderately correlated with each other and are associated with different known predictors of BC. The classification of genetic risk for the same individual may vary considerably depending on the chosen GRS.\n\nConclusionsWe have shown that metaGRS2 that combines on the effects of more than 900 SNPs provides best predictive ability for breast cancer in two different population-based cohorts. The strength of the effect of metaGRS2 indicates that the GRS could potentially be used to develop more efficient strategies for breast cancer screening for genotyped women.

genetics

The evolution of genetic diversity in changing environments

The production and maintenance of genetic and phenotypic diversity under temporally fluctuating selection and the signatures of environmental and selective volatility in the patterns of genetic and phenotypic variation have been important areas of focus in population genetics. On one hand, stretches of constant selection pull the genetic makeup of populations towards local fitness optima. On the other, in order to cope with changes in the selection regime, populations may evolve mechanisms that create a diversity of genotypes. By tuning the rates at which variability is produced, such as the rates of recombination, mutation or migration, populations may increase their long-term adaptability. Here we use theoretical models to gain insight into how the rates of these three evolutionary forces are shaped by fluctuating selection. We compare and contrast the evolution of recombination, mutation and migration under similar patterns of environmental change and show that these three sources of phenotypic variation are surprisingly similar in their response to changing selection. We show that knowing the shape, size, variance and asymmetry of environmental runs is essential for accurate prediction of genetic evolutionary dynamics.

Evolutionary Biology

THE GENETIC LANDSCAPE OF TRANSCRIPTIONAL NETWORKS IN A COMBINED HAPLOID/DIPLOID PLANT SYSTEM

Heritable variation in gene expression is a source of evolutionary change and our understanding of the genetic basis of expression variation remains incomplete. Here, we dissected the genetic basis of transcriptional variation in a wild, outbreeding gymnosperm (Picea glauca) according to linked and unlinked genetic variants, their allele-specific (cis) and allele non-specific (trans) effects, and their phenotypic additivity. We used a novel plant system that is based on the analysis of segregating alleles of a single self-fertilized plant in haploid and diploid seed tissues. We measured transcript abundance and identified transcribed SNPs in 66 seeds with RNA-seq. Linked and unlinked genetic effects that influenced expression levels were abundant in the haploid megagametophyte tissue, influencing 48% and 38% of analyzed genes, respectively. Analysis of these effects in diploid embryos revealed that while distant effects were acting in trans consistent with their hypothesized diffusible nature, local effects were associated with a complex mix of cis, trans and compensatory effects. Most cis effects were additive irrespective of their effect sizes, consistent with a hypothesis that they represent rate-limiting factors in transcript accumulation. We show that trans effects fulfilled a key prediction of Wright s physiological theory, in which variants with small effects tend to be additive and those with large effects tend to be dominant/recessive. Our haploid/diploid approach allows a comprehensive genetic dissection of expression variation and can be applied to a large number of wild plant species.

Genomics