Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Evolutionary design of regulatory control. II. Robust error-correcting feedback increases genetic and phenotypic variability

As systems become more robust against perturbations, they can compensate for greater sloppiness in the performance of their components. That robust compensation reduces the force of natural selection on the systems components, leading to component decay. The paradoxical coupling of robustness and decay predicts that robust systems evolve cheaper, lower performing components, which accumulate greater mutational genetic variability and which have greater phenotypic stochasticity in trait expression. Previous work noted the paradox of robustness. However, no general theory for the evolutionary dynamics of system robustness and component decay has been developed. This article takes a first step by linking engineering control theory with the genetic theory of evolutionary dynamics. Control theory emphasizes error-correcting feedback as the single greatest principle in robust system design. Linking control theory to evolution leads to a theory for the evolutionary dynamics of error-correcting feedback, a unifying approach for the evolutionary analysis of robust systems. In this article, I study how, in theory, increasingly robust systems accumulate more genetic variability and greater stochasticity of expression in their components. The theory predicts different levels of variability between different regulatory control architectures and different levels of variability between different components within a particular regulatory control system. Those predictions provide a way to understand the accumulating data on genetic variability and single-cell stochasticity of gene expression. I also show that increasing robustness reduces the frequency of system failures associated with disease and, simultaneously, causes a strong increase in the heritability of disease. Thus, robust error correction in biological regulatory control may partly explain the puzzlingly high heritability of disease and, more generally, the surprisingly high heritability of fitness.

systems biology

A General Framework for Variable Selection in Linear Mixed Models with Applications to Genetic Studies with Structured Populations

Complex traits are known to be influenced by a combination of environmental factors and rare and common genetic variants. However, detection of such multivariate associations can be compromised by low statistical power and confounding by population structure. Linear mixed effects models (LMM) can account for correlations due to relatedness but have not been applicable in high-dimensional (HD) settings where the number of fixed effect predictors greatly exceeds the number of samples. False positives or false negatives can result from two-stage approaches, where the residuals estimated from a null model adjusted for the subjects relationship structure are sub-sequently used as the response in a standard penalized regression model. To overcome these challenges, we develop a general penalized LMM with a single random effect called ggmix for simultaneous SNP selection and adjustment for population structure in high dimensional prediction models. We develop a blockwise coordinate descent algorithm with automatic tuning parameter selection which is highly scalable, computationally efficient and has theoretical guarantees of convergence. Through simulations and three real data examples, we show that ggmix leads to more parsimonious models compared to the two-stage approach or principal component adjustment with better prediction accuracy. Our method performs well even in the presence of highly correlated markers, and when the causal SNPs are included in the kinship matrix. ggmix can be used to construct polygenic risk scores and select instrumental variables in Mendelian randomization studies. Our algorithms are available in an R package (https://github.com/greenwoodlab/ggmix). 1 Author SummaryThis work addresses a recurring challenge in the analysis and interpretation of genetic association studies: which genetic variants can best predict and are independently associated with a given phenotype in the presence of population structure ? Not controlling confounding due to geographic population structure, family and/or cryptic relatedness can lead to spurious associations. Much of the existing research has therefore focused on modeling the association between a phenotype and a single genetic variant in a linear mixed model with a random effect. However, this univariate approach may miss true associations due to the stringent significance thresholds required to reduce the number of false positives and also ignores the correlations between markers. We propose an alternative method for fitting high-dimensional multivariable models, which selects SNPs that are independently associated with the phenotype while also accounting for population structure. We provide an efficient implementation of our algorithm and show through simulation studies and real data examples that our method outperforms existing methods in terms of prediction accuracy and controlling the false discovery rate.

bioinformatics

Construction of Feed Forward MultiLayer Perceptron Model For Genetic Dataset in Leishmaniasis Using Cognitive Computing

Leishmaniasis is an endemic parasitic disease, predominantly found in the poor locality of Africa, Asia and Latin America. It is associated with malnutrition, weak immune system of people and their housing locality. At present, it is diagnosed by microscopic identification, molecular and biochemical characterisation or serum analysis for parasitic compounds. In this study, we present a new approach for diagnosing Leishmaniasis using cognitive computing. The Genetic datasets of leishmaniasis are collected from Gene Expression Omnibus database and its then processed. The algorithm for training and developing a model, based on the data is prepared and coded using python. The algorithm and their corresponding datasets are integrated using TensorFlow dataframe. A feed forward Artificial Neural Network trained model with multi-layer perceptron is developed as a diagnosing model for Leishmaniasis, using genetic dataset. It is developed using recurrent neural network. The cognitive model of the trained network is interpreted using the maps and mathematical formula of the influencing parameters. The credit of the system is measured using the accuracy, loss and error of the system. This integrated system of the leishmaniasis genetic dataset and neural network proved to be the good choice for diagnosis with higher accuracy and lower error. Through this approach, all records of the data are effectively incorporated into the system. The experimental results of feed forward multilayer perceptron model after normalization; mean square error (219.84), loss function (1.94) and accuracy (85.71%) of the model, shows good fit of model with the process and it could possibly serve as a better solution for diagnosing Leishmaniasis in future, using genetic datasets.\n\nThe code is available in Github repository:\n\nhttps://github.com/shailzasingh/Machine-Learning-code-for-analyzing-genetic-dataset-in-Leishmaniasis

bioinformatics

Pleistocene climate changes explain large-scale genetic variation in a dominant grassland species, Lolium perenne L.

AimGrasslands have been pivotal in the development of herbivore breeding since the Neolithic and are still nowadays the most widespread agricultural land-use across Europe. However, it remains unclear whether the current large-scale genetic variation of plant species found in natural grasslands of Europe is the result of human activities or natural processes.\n\nLocationEurope.\n\nTaxonLolium perenne L (perennial ryegrass).\n\nMethodsWe reconstructed the phylogeographic history of L. perenne, a dominant grassland species, using 481 natural populations including 11 populations from closely related taxa. We combined the Genotyping-by-Sequencing (GBS) and Pool-sequencing (Pool-seq) methods to obtain high-quality allele frequency calls of ~ 500 k SNP loci. We performed genetic structure analyses and demographic reconstructions based on the site frequency spectrum (SFS). We additionally used the same genotyping protocol to assess the genomic diversity of a set of 32 cultivars representative of the L. perenne cultivars widely used for forage purposes.\n\nResultsExpansion across Europe took place during the Wurm glaciation (12-110 kya), a cooling period that decreased the dominance of trees in favour of grasses. Splits and admixtures in L. perenne fit historical sea level changes in the Mediterranean basin. The development of agriculture in Europe (7-3.5 kya), that caused an increase in the abundance of grasslands, did not have an effect on the demographic patterns of L. perenne. We found little differentiation between modern cultivars and certain natural variants. However, modern cultivars do not represent the wide genetic variation found in natural populations.\n\nMain conclusionsDemographic events in L. perenne can be explained by the changing climatic conditions during the Pleistocene. Natural populations maintain a wide genomic variability at continental scale that has been underused by recent breeding activities. This variability constitutes valuable standing genetic variation for future adaptation of grasslands to climate change, safeguarding the agricultural services they provide.

evolutionary biology

Population structure, genetic connectivity, and adaptation in the Olympia oyster (Ostrea lurida) along the west coast of North America

Effective management of threatened and exploited species requires an understanding of both the genetic connectivity among populations and local adaptation. The Olympia oyster (Ostrea lurida), patchily distributed from Baja California to the central coast of Canada, has a long history of population declines due to anthropogenic stressors. For such coastal marine species, population structure could follow a continuous isolation-by-distance model, contain regional blocks of genetic similarity separated by barriers to gene flow, or be consistent with a null model of no population structure. To distinguish between these hypotheses in O. lurida, 13,444 single-nucleotide polymorphisms (SNPs) were used to characterize rangewide population structure, genetic connectivity, and adaptive divergence. Samples were collected across the species range on the west coast of North America, from southern California to Vancouver Island. A conservative approach for detecting putative loci under selection identified 288 SNPs across 129 GBS loci, which were functionally annotated and analyzed separately from the remaining neutral loci. While strong population structure was observed on a regional scale in both neutral and outlier markers, neutral markers had greater power to detect fine-scale structure. Geographic regions of reduced gene flow aligned with known marine biogeographic barriers, such as Cape Mendocino, Monterey Bay, and the currents around Cape Flattery. The outlier loci identified as under putative selection included genes involved in developmental regulation, sensory information processing, energy metabolism, immune response, and muscle contraction. These loci are excellent candidates for future research and may provide targets for genetic monitoring programs. Beyond specific applications for restoration and management of the Olympia oyster, this study lends to the growing body of evidence for both population structure and adaptive differentiation across a range of marine species exhibiting the potential for panmixia. Computational notebooks are available to facilitate reproducibility and future open-sourced research on the population structure of O. lurida.

evolutionary biology

SLiM 3: Forward genetic simulations beyond the Wright-Fisher model

With the desire to model population genetic processes under increasingly realistic scenarios, forward genetic simulations have become a critical part of the toolbox of modern evolutionary biology. The SLiM forward genetic simulation framework is one of the most powerful and widely used tools in this area. However, its foundation in the Wright-Fisher model has been found to pose an obstacle to implementing many types of models; it is difficult to adapt the Wright-Fisher model, with its many assumptions, to modeling ecologically realistic scenarios such as explicit space, overlapping generations, individual variation in reproduction, density-dependent population regulation, individual variation in dispersal or migration, local extinction and recolonization, mating between subpopulations, age structure, fitness-based survival and hard selection, emergent sex ratios, and so forth. In response to this need, we here introduce SLiM 3, which contains two key advancements aimed at abolishing these limitations. First, the new non-Wright-Fisher or \"nonWF\" model type provides a much more flexible foundation that allows the easy implementation of all of the above scenarios and many more. Second, SLiM 3 adds support for continuous space, including spatial interactions and spatial maps of environmental variables. We provide a conceptual overview of these new features, and present several example models to illustrate their use. These two key features allow SLiM 3 models to go beyond the Wright-Fisher model, opening up new horizons for forward genetic modeling.

evolutionary biology

Rare genetic variation is important for survival in extreme low pH conditions

Standing genetic variation is important for population persistence in extreme environmental conditions. While some species may have the capacity to adapt to predicted average future global change conditions, the ability to survive extreme events is largely unknown. We used single generation selection experiments on hundreds of thousands of Strongylocentrotus purpuratus sea urchin larvae generated from wild-caught adults to identify adaptive genetic variation responsive to moderate (pH 8.0) and extreme (pH 7.5) low pH conditions. Sequencing genomic DNA from pools of larvae, we identified consistent changes in allele frequencies across replicate cultures of both conditions and observed increased linkage disequilibrium around selected loci, revealing selection on recombined standing genetic variation. We found that loci responding uniquely to either selection regime were at low starting allele frequencies while variants that responded to both pH conditions (11.6% of selected variants) started at high frequencies. Loci under selection performed functions related to energetics, pH tolerance, cell growth, and actin/cytoskeleton dynamics. These results highlight that persistence in future conditions will require two classes of genetic variation: common, pH-responsive variants maintained by balancing selection in a heterogeneous environment, and rare variants, particularly for extreme conditions, that must be maintained by large population sizes.

evolutionary biology

Genetic basis and timing of a major mating system shift in Capsella

Shifts from outcrossing to self-fertilisation have occurred repeatedly in many different lineages of flowering plants, and often involve the breakdown of genetic outcrossing mechanisms. In the Brassicaceae, self-incompatibility (SI) allows plants to ensure outcrossing by recognition and rejection of self-pollen on the stigma. This occurs through the interaction of female and male specificity components, consisting of a pistil based receptor and a pollen-coat protein, both of which are encoded by tightly linked genes at the S-locus. When benefits of selfing are higher than costs of inbreeding, theory predicts that loss-of-function mutations in the male (pollen) SI component should be favoured, especially if they are dominant. However, it remains unclear whether mutations in the male component of SI are predominantly responsible for shifts to self-compatibility, and testing this prediction has been difficult due to the challenges of sequencing the highly polymorphic and repetitive ~100 kbp S-locus. The crucifer genus Capsella offers an excellent opportunity to study multiple transitions from outcrossing to self-fertilization, but so far, little is known about the genetic basis and timing of loss of SI in the self-fertilizing diploid Capsella orientalis. Here, we show that loss of SI in C. orientalis occurred within the past 2.6 Mya and maps as a dominant trait to the S-locus. Using targeted long-read sequencing of multiple complete S-haplotypes, we identify a frameshift deletion in the male specificity gene SCR that is fixed in C. orientalis, and we confirm loss of male SI specificity. We further analyze RNA sequencing data to identify a conserved, S-linked small RNA (sRNA) that is predicted to cause dominance of self-compatibility. Our results suggest that degeneration of pollen SI specificity in dominant S-alleles is important for shifts to self-fertilization in the Brassicaceae.\n\nAuthor SummaryAlready Darwin was fascinated by the widely varying modes of plant reproduction. The shift from outcrossing to self-fertilization is considered one of the most frequent evolutionary transitions in flowering plants, yet we still know little about the genetic basis of these shifts. In the Brassicaceae, outcrossing is enforced by a self-incompatibility (SI) system that enables the recognition and rejection of self pollen. This occurs through the action of two tightly linked genes at the S-locus, that encode a receptor protein located on the stigma (female component) and a pollen ligand protein (male component), respectively. Nevertheless, SI has frequently been lost, and theory predicts that mutations in the male component should have an advantage during the loss of SI, especially if they are dominant. To test this hypothesis, we mapped the loss of SI in a selfing species from the genus Capsella, a model system for evolutionary genomics. We found that loss of SI mapped to the S-locus, which harbored a dominant loss-of-function mutation in the male SI protein, and as expected, we found that male specificity was indeed lost in C. orientalis. Our results suggest that transitions to selfing often involve parallel genetic changes.

evolutionary biology

Genetic variants influence on the placenta regulatory landscape

BackgroundFrom genomic association studies, quantitative trait loci analysis, and epigenomic mapping, it is evident that significant efforts are necessary to define genetic-epigenetic interactions and understand their role in disease susceptibility and progression. For this reason, an analysis of the effects of genetic variation on gene expression and DNA methylation in human placentas at high resolution and whole-genome coverage will have multiple mechanistic and practical implications.\n\nResultsBy producing and analyzing DNA sequence variation (n=303), DNA methylation (n=303) and mRNA expression data (n=80) from placentas from healthy women, we investigate the regulatory landscape of the human placenta and offer analytical approaches to integrate different types of genomic data and address some potential limitations of current platforms. We distinguish two profiles of interaction between expression and DNA methylation, revealing linear or bimodal effects, reflecting differences in genomic context, transcription factor recruitment, and possibly cell subpopulations.\n\nConclusionsThese findings help to clarify the interactions of genetic, epigenetic, and transcriptional regulatory mechanisms in normal human placentas. They also provide strong evidence for genotype-driven modifications of transcription and DNA methylation in normal placentas. In addition to these mechanistic implications, the data and analytical methods presented here will improve the interpretability of genome-wide and epigenome-wide association studies for human traits and diseases that involve placental functions.\n\nAuthor summaryThe placenta is a critical organ playing multiple roles including oxygen and metabolite transfer from mother to fetus, hormone production, and vascular perfusion. With this study, we aimed to deliver a placenta-specific regulatory map based on a combination of publicly available and newly generated data. To complete this reference, we obtained genotype information (n=303), DNA methylation (n=303) and expression data (n=80) for placentas from healthy women. Our analysis of methylation and expression quantitative trait loci (QTLs) and correlations between methylation and expression data were designed to identify fundamental associations between genome, transcriptome, and epigenome in this key fetal organ. The results provide high-resolution genetic and epigenetic maps specific to the placenta based on a representative ethnically diverse cohort. As interest and efforts are growing to better understand the etiology of placental disease and the impact of the environment on placental function these data will provide a reference and enhance future investigations.

genomics

Genetic control of variability in subcortical and intracranial volumes

Sensitivity to external demands is essential for adaptation to dynamic environments, but comes at the cost of increased risk of adverse outcomes when facing poor environmental conditions. Here, we apply a novel methodology to perform genome-wide association analysis of mean and variance in nine key brain features (accumbens, amygdala, caudate, hippocampus, pallidum, putamen, thalamus, intracranial volume and cortical thickness), integrating genetic and neuroanatomical data from a large lifespan sample (n=25,575 individuals; 8 to 89 years, mean age 51.9 years). We identify genetic loci associated with phenotypic variability in cortical thickness, thalamus, pallidum, and intracranial volumes. The variance-controlling loci included genes with a documented role in brain and mental health and were not associated with the mean anatomical volumes. This proof-of-principle of the hypothesis of a genetic regulation of brain volume variability contributes to establishing the genetic basis of phenotypic variance (i.e., heritability), allows identifying different degrees of brain robustness across individuals, and opens new research avenues in the search for mechanisms controlling brain and mental health.

neuroscience

Genetic distance and social compatibility in the aggregation behavior of Japanese toad tadpoles

From microorganism to vertebrates, living things often exhibit social aggregation. One of anuran larvae, dark-bodied toad tadpoles (genus Bufo) are known to aggregate against predators. When individuals share genes from a common ancestor for whom social aggregation was a functional trait, they are also likely to share common recognition cues regarding association preferences, while greater genetic distances make cohesive aggregation difficult. In this study, we conducted quantitative analyses to examine aggregation behavior among three lineages of toad tadpoles: Bufo japonicus japonicus, B. japonicus formosus, and B. gargarizans miyakonis. To determine whether there is a correlation between cohesiveness and genetic similarity among group members, we conducted an aggregation test using 42 cohorts consisting of combinations drawn from a laboratory-reared set belonging to distinct clutches. As genetic indices, we used mitochondrial DNA (mtDNA) and major histocompatibility complex (MHC) class II alleles. The results clearly indicated that aggregation behavior in toad tadpoles is directly influenced by genetic distances based on mtDNA sequences and not on MHC haplotypes. Cohesiveness among heterogeneous tadpoles is negatively correlated with the geographic dispersal of groups. Our findings suggest that social incompatibility among toad tadpoles reflects phylogenetic relationships.

animal behavior and cognition

Genetic Associations with Subjective Well-Being Also Implicate Depression and Neuroticism

We conducted a genome-wide association study of subjective well-being (SWB) in 298,420 individuals. We also performed auxiliary analyses of depressive symptoms (\"DS\"; N = 161,460) and neuroticism (N = 170,910), both of which have a substantial genetic correlation with SWB [Formula]. We identify three SNPs associated with SWB at genome-wide significance. Two of them are significantly associated with DS in an independent sample. In our auxiliary analyses, we identify 13 additional genome-wide-significant associations: two with DS and eleven with neuroticism, including two inversion polymorphisms. Across our phenotypes, loci regulating expression in central nervous system and adrenal/pancreas tissues are enriched. The discovery of genetic loci associated with the three phenotypes we study has proven elusive; our findings illustrate the payoffs from studying them jointly.\n\nOne Sentence Summary: Using both genome-wide association studies and proxy-phenotype studies, we identify genetic variants associated with subjective well-being, depressive symptoms, and neuroticism.

Genetics

Invasion genetics of the silver carp (Hypophthalmichthys molitrix) across North America: Differentiation of fronts, introgression, and eDNA detection

The invasive silver carp Hypophthalmichthys molitrix escaped from southern U.S. aquaculture during the 1970s to spread throughout the Mississippi River basin and steadily moved northward, now reaching the threshold of the Laurentian Great Lakes. The silver carp is native to eastern Asia and is a large, prolific filter-feeder that decreases food availability for fisheries. The present study evaluates its population genetic variability and differentiation across the introduced range using 10 nuclear DNA microsatellite loci, sequences of two mitochondrial genes (cytochrome b and cytochrome c oxidase subunit 1), and a nuclear gene (ribosomal protein S7 gene intron 1). Populations are analyzed from two invasion fronts threatening the Great Lakes (the Illinois River outside Lake Michigan and the Wabash River, leading into the Maumee River and western Lake Erie), established areas in the southern and central Mississippi River, and a later Missouri River colonization. Results discern considerable genetic diversity and some significant population differentiation, with greater mtDNA haplotype diversity and unique microsatellite alleles characterizing the southern populations. Invasion fronts significantly differ, diverging from the southern Mississippi River population. About 3% of individuals contain a unique and very divergent mtDNA haplotype (primarily the southerly populations and the Wabash River), which may stem from historic introgression in Asia with female largescale silver carp H. harmandi. Nuclear microsatellites and S7 sequences of the introgressed individuals do not significantly differ from silver carp. MtDNA variation is used in a high-throughput sequence assay that identifies and distinguishes invasive carp species and their population haplotypes (including H. molitrix and H. harmandi) at all life stages, in application to environmental (e)DNA water and plankton samples. We discerned silver and bighead carp eDNA from four bait and pond stores in the Great Lakes watershed, indicating that release from retailers comprises another likely vector. Our findings provide key baseline population genetic data for understanding and tracing the invasions progression, facilitating detection, and evaluating future trajectory and adaptive success.

genetics

Estimating strength of polygenic selection with principal components analysis of spatial genetic variation

Principal components analysis on allele frequencies for 14 and 50 populations (from 1K Genomes and ALFRED databases) produced a factor accounting for over half of the variance, which indicates selection pressure on intelligence or genotypic IQ. Very high correlations between this factor and phenotypic IQ, educational achievement were observed (r>0.9 and r>0.8), also after partialling out GDP and the Human Development Index. Regression analysis was used to estimate a genotypic (predicted) IQ also for populations with missing data for phenotypic IQ. Socio-economic indicators (GDP and Human Development Index) failed to predict residuals, not providing evidence for the effects of environmental factors on intelligence. Another analysis revealed that the relationship between IQ and the genotypic factor was not mediated by race, implying that it exists at a finer resolution, a finding which in turn suggests selective pressures postdating sub-continental population splits.\n\nGenotypic height and IQ were inversely correlated but this correlation was mostly mediated by race. In at least two cases (Native Americans vs East Asians and Africans vs Papuans) genetic distance inferred from evolutionarily neutral genetic markers contrasts markedly with the resemblance observed for IQ and height increasing alleles.\n\nA principal component analysis on a random sample of 20 SNPs revealed two factors representing genetic relatedness due to migrations. However, the correlation between IQ and the intelligence PC was not mediated by them. In fact, the intelligence PC emerged as an even stronger predictor of IQ after entering the \"migratory\" PCs in a regression, indicating that it represents selection pressure instead of migrational effects.\n\nFinally, some observations on the high IQ of Mongoloid people are made which lend support to the \"cold winters theory\" on the evolution of intelligence.

Genetics

The genetic ancestry of African, Latino, and European Americans across the United States.

Over the past 500 years, North America has been the site of ongoing mixing of Native Americans, European settlers, and Africans brought largely by the Trans-Atlantic slave trade, shaping the early history of what became the United States. We studied the genetic ancestry of 5,269 self-described African Americans, 8,663 Latinos, and 148,789 European Americans who are 23andMe customers and show that the legacy of these historical interactions is visible in the genetic ancestry of present-day Americans. We document pervasive mixed ancestry and asymmetrical male and female ancestry contributions in all groups studied. We show that regional ancestry differences reflect historical events, such as early Spanish colonization, waves of immigration from many regions of Europe, and forced relocation of Native Americans within the US. This study sheds light on the fine-scale differences in ancestry within and across the United States, and informs our understanding of the relationship between racial and ethnic identities and genetic ancestry.

Genetics

Genetic interactions contribute less than additive effects to quantitative trait variation in yeast

Genetic mapping studies of quantitative traits typically focus on detecting loci that contribute additively to trait variation. Genetic interactions are often proposed as a contributing factor to trait variation, but the relative contribution of interactions to trait variation is a subject of debate. Here, we use a very large cross between two yeast strains to accurately estimate the fraction of phenotypic variance due to pairwise QTL-QTL interactions for 20 quantitative traits. We find that this fraction is 9% on average, substantially less than the contribution of additive QTL (43%). Statistically significant QTL-QTL pairs typically have small individual effect sizes, but collectively explain 40% of the pairwise interaction variance. We show that pairwise interaction variance is largely explained by pairs of loci at least one of which has a significant additive effect. These results refine our understanding of the genetic architecture of quantitative traits and help guide future mapping studies.

Genetics

The anatomical distribution of genetic associations

Deeper understanding of the anatomical intermediaries for disease and other complex genetic traits is essential to understanding mechanisms and developing new interventions. Existing ontology tools provide functional annotations for many genes in the genome and they are widely used to develop mechanistic hypotheses based on genetic and transcriptomic data. Yet, information about where a set of genes is expressed may be equally useful in interpreting results and forming novel mechanistic hypotheses for a trait. Therefore, we developed a framework for statistically testing the relationship between gene expression across the body and sets of candidate genes from across the genome. We validated this tool and tested its utility on three applications. First, using thousands of loci identified by GWA studies, our framework identifies the number of disease-associated genes that have enriched expression in the disease-affected tissue. Second, we experimentally confirmed an underappreciated prediction highlighted by our tool: variation in skin expressed genes are a major quantitative genetic modulator of white blood cell count - a trait considered to be a feature of the immune system. Finally, using gene lists derived from sequencing data, we show that human genes under constrained selective pressure are disproportionately expressed in nervous system tissues.

Genetics

The effects of both recent and long-term selection and genetic drift are readily evident in North American barley breeding populations

Barley was introduced to North America [~]400 years ago but adaptation to modern production environments is more recent. Comparisons of allele frequencies among different growth habits and inflorescence types in North America indicate significant genetic differentiation has accumulated in a relatively short evolutionary time span. Allele frequency differentiation is greatest among barley with two-row versus six-row inflorescences, and then by spring versus winter growth habit. Large changes in allele frequency among breeding programs suggest a major contribution of genetic drift and linked selection on genetic variation. Despite this, comparisons of 3,613 modern North American cultivated breeding lines that differ for row type and growth habit permit the discovery of 183 SNP outliers putatively linked to targets of selection. For example, SNPs within the Cbf4, Ppd-H1, and Vrn-H1 loci which have previously been associated with agronomically-adaptive phenotypes, are identified as outliers. Analysis of extended haplotype-sharing identifies genomic regions shared within and among breeding programs, suggestive of a number of genomic regions subject to recent selection. Finally, we are able to identify recent bouts of gene flow between breeding programs that could point to the sharing of agronomically-adaptive variation. These results are supported by pedigrees and breeders understanding of germplasm sharing.

Genetics