Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The maternal genetic make-up of the Iberian Peninsula between the Neolithic and the Early Bronze Age

Agriculture first reached the Iberian Peninsula around 5700 BCE. However, little is known about the genetic structure and changes of prehistoric populations in different geographic areas of Iberia. In our study, we focused on the maternal genetic makeup of the Neolithic ([~] 5500-3000 BCE), Chalcolithic ([~] 3000-2200 BCE) and Early Bronze Age ([~] 2200-1500 BCE). We report ancient mitochondrial DNA results of 213 individuals (151 HVS-I sequences) from the northeast, central, southeast and southwest regions and thus on the largest archaeogenetic dataset from the Peninsula to date. Similar to other parts of Europe, we observe a discontinuity between hunter-gatherers and the first farmers of the Neolithic. During the subsequent periods, we detect regional continuity of Early Neolithic lineages across Iberia, however the genetic contribution of hunter-gatherers is generally higher than in other parts of Europe and varies regionally. In contrast to ancient DNA findings from Central Europe, we do not observe a major turnover in the mtDNA record of the Iberian Late Chalcolithic and Early Bronze Age, suggesting that the population history of the Iberian Peninsula is distinct in character.

genetics

Genetics of educational attainment aid in identifying biological subcategories of schizophrenia

Higher educational attainment (EA) is negatively associated with schizophrenia (SZ). However, recent studies found a positive genetic correlation between EA and SZ. We investigated possible causes of this counterintuitive finding using genome-wide association study results for EA and SZ (N = 443,581) and a replication cohort (1,169 controls; 1,067 cases) with deeply phenotyped SZ patients. We found strong genetic dependence between EA and SZ that cannot be explained by chance, linkage disequilibrium, or assortative mating. Instead, several genes seem to have pleiotropic effects on EA and SZ, but without a clear pattern of sign concordance. Genetic heterogeneity of SZ contributes to this finding. We demonstrate this by showing that the polygenic prediction of clinical SZ symptoms can be improved by taking the sign concordance of loci for EA and SZ into account. Furthermore, using EA as a proxy phenotype, we isolate FOXO6 and SLITRK1 as novel candidate genes for SZ.

genetics

Application of t-SNE to Human Genetic Data

The t-SNE (t-distributed stochastic neighbor embedding) is a new dimension reduction and visualization technique for high-dimensional data. t-SNE is rarely applied to human genetic data, even though it is commonly used in other data-intensive biological fields, such as single-cell genomics. We explore the applicability of t-SNE to human genetic data and make these observations: (i) similar to previously used dimension reduction techniques such as principal component analysis (PCA), t-SNE is able to separate samples from different continents; (ii) unlike PCA, t-SNE is more robust with respect to the presence of outliers; (iii) t-SNE is able to display both continental and sub-continental patterns in a single plot. We conclude that the ability for t-SNE to reveal population stratification at different scales could be useful for human genetic association studies.

genetics

Genetic regulatory effects modified by immune activation contribute to autoimmune disease associations

The immune system plays a major role in human health and disease, and understanding genetic causes of interindividual variability of immune responses is vital. We isolated monocytes from 134 genotyped individuals, stimulated the cells with three defined microbe-associated molecular patterns (LPS, MDP, and ppp-dsRNA), and profiled the transcriptome at three time points. After mapping expression quantitative trait loci (eQTL), we identified 417 response eQTLs (reQTLs) with differing effect between the conditions. We characterized the dynamics of genetic regulation on early and late immune response, and observed an enrichment of reQTLs in distal cis-regulatory elements. Response eQTLs are also enriched for recent positive selection with an evolutionary trend towards enhanced immune response. Finally, we uncover novel reQTL effects in multiple GWAS loci, and show a stronger enrichment of response than constant eQTLs in GWAS signals of several autoimmune diseases. This demonstrates the importance of infectious stimuli modifying genetic predisposition to disease.

genetics

Genome-wide association study of alcohol consumption and genetic overlap with other health-related traits in UK Biobank (N=112,117).

Alcohol consumption has been linked to over 200 diseases and is responsible for over 5% of the global disease burden. Well known genetic variants in alcohol metabolizing genes, e.g. ALDH2, ADH1B, are strongly associated with alcohol consumption but have limited impact in European populations where they are found at low frequency. We performed a genome-wide association study (GWAS) of self-reported alcohol consumption in 112,117 individuals in the UK Biobank (UKB) sample of white British individuals. We report significant genome-wide associations at 8 independent loci. These include SNPs in alcohol metabolizing genes (ADH1B/ADH1C/ADH5) and 2 loci in KLB, a gene recently associated with alcohol consumption. We also identify SNPs at novel loci including GCKR, PXDN, CADM2 and TNFRSF11A. Gene-based analyses found significant associations with genes implicated in the neurobiology of substance use (CRHR1, DRD2), and genes previously associated with alcohol consumption (AUTS2). GCTA-GREML analyses found a significant SNP-based heritability of self-reported alcohol consumption of 13% (S.E.=0.01). Sex-specific analyses found largely overlapping GWAS loci and the genetic correlation between male and female alcohol consumption was 0.73 (S.E.=0.09, p-value = 1.37 x 10-16). Using LD score regression, genetic overlap was found between alcohol consumption and schizophrenia (rG=0.13, S.E=0.04), HDL cholesterol (rG=0.21, S.E=0.05), smoking (rG=0.49, S.E=0.06) and various anthropometric traits (e.g. Overweight, rG=-0.19, S.E.=0.05). This study replicates the association between alcohol consumption and alcohol metabolizing genes and KLB, and identifies 4 novel gene associations that should be the focus of future studies investigating the neurobiology of alcohol consumption.

genetics

Contribution Of Trans Regulatory eQTL To Cryptic Genetic Variation In C. elegans

BackgroundCryptic genetic variation (CGV) is the hidden genetic variation that can be unlocked by perturbing normal conditions. CGV can drive the emergence of novel complex phenotypes through changes in gene expression. Although our theoretical understanding of CGV has thoroughly increased over the past decade, insight into polymorphic gene expression regulation underlying CGV is scarce. Here we investigated the transcriptional architecture of CGV in response to rapid temperature changes in the nematode Caenorhabditis elegans. We analyzed regulatory variation in gene expression (and mapped eQTL) across the course of a heat stress and recovery response in a recombinant inbred population.\n\nResultsWe measured gene expression over three temperature treatments: i) control, ii) heat stress, and iii) recovery from heat stress. Compared to control, exposure to heat stress affected the transcription of 3305 genes, whereas 942 were affected in recovering worms. These affected genes were mainly involved in metabolism and reproduction. The gene expression pattern in recovering worms resembled both the control and the heat stress treatment. We mapped eQTL using the genetic variation of the recombinant inbred population and detected 2626 genes with an eQTL in the heat stress treatment, 1797 in the control, and 1880 in the recovery. The cis-eQTL were highly conserved across treatments. A considerable fraction of the trans-eQTL (40-57%) mapped to 19 treatment specific trans-bands. In contrast to cis-eQTL, trans-eQTL were highly environment specific and thus cryptic. Approximately 67% of the trans-eQTL were only induced in a single treatment, with heat-stress showing the most unique trans-eQTL.\n\nConclusionsThese results illustrate the highly dynamic pattern of CGV across three different environmental conditions that can be evoked by a stress response over a relatively short time-span (2 hours) and that CGV is mainly determined by response related trans regulatory eQTL.

genetics

Extending Causality Tests With Genetic Instruments: An Integration Of Mendelian Randomization And The Classical Twin Design

Mendelian Randomization (MR) is an important approach to modelling causality in non-experimental settings. MR uses genetic instruments to test causal relationships between exposures and outcomes of interest. Individual genetic variants have small effects, and so, when used as instruments, render MR liable to weak instrument bias. Polygenic scores have the advantage of larger effects, but may be characterized by direct pleiotropy, which violates a central assumption of MR.\n\nWe developed the MR-DoC twin model by integrating MR with the Direction of Causation twin model. This model allows us to test pleiotropy directly. We considered the issue of parameter identification, and given identification, we conducted extensive power calculations. MR-DoC allows one to test causal hypotheses and to obtain unbiased estimates of the causal effect given pleiotropic instruments (polygenic scores), while controlling for genetic and environmental influences common to the outcome and exposure. Furthermore, MR-DoC in twins has appreciably greater statistical power than a standard MR analysis applied to singletons, if the unshared environmental effects on the exposure and the outcome are uncorrelated. Generally, power increases with: 1) decreasing residual exposure-outcome correlation, and 2) decreasing heritability of the exposure variable.\n\nMR-DoC allows one to employ strong instrumental variables (polygenic scores, possibly pleiotropic), guarding against weak instrument bias and increasing the power to detect causal effects. Our approach will enhance and extend MRs range of applications, and increase the value of the large cohorts collected at twin registries as they correctly detect causation and estimate effect sizes even in the presence of pleiotropy.

genetics

lme4qtl: Linear Mixed Models With Flexible Covariance Structure For Genetic Studies Of Related Individuals

BackgroundQuantitative trait locus (QTL) mapping in genetic data often involves analysis of correlated observations, which need to be accounted for to avoid false association signals. This is commonly performed by modeling such correlations as random effects in linear mixed models (LMMs). The R package lme4 is a well-established tool that implements major LMM features using sparse matrix methods; however, it is not fully adapted for QTL mapping association and linkage studies. In particular, two LMM features are lacking in the base version of lme4: the definition of random effects by custom covariance matrices; and parameter constraints, which are essential in advanced QTL models. Apart from applications in linkage studies of related individuals, such functionalities are of high interest for association studies in situations where multiple covariance matrices need to be modeled, a scenario not covered by many genome-wide association study (GWAS) software.\n\nResultsTo address the aforementioned limitations, we developed a new R package lme4qtl as an extension of lme4. First, lme4qtl contributes new models for genetic studies within a single tool integrated with lme4 and its companion packages. Second, lme4qtl offers a flexible framework for scenarios with multiple levels of relatedness and becomes efficient when covariance matrices are sparse. We showed the value of our package using real family-based data in the Genetic Analysis of Idiopathic Thrombophilia 2 (GAIT2) project.\n\nConclusionsOur software lme4qtl enables QTL mapping models with a versatile structure of random effects and efficient computation for sparse covariances. lme4qtl is available at https://github.com/variani/lme4qtl.

genetics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics

Genetic contribution to two factors of neuroticism is associated with affluence, better health, and longer life

Neuroticism is a personality trait that describes the tendency to experience negative emotions. Individual differences in neuroticism are moderately stable across much of the life course1; the trait is heritable2-5, and higher levels are associated with psychiatric disorders6-8, and have been estimated to have an economic burden to society greater than that of substance abuse, mood, or anxiety disorders9. Understanding the genetic architecture of neuroticism therefore has the potential to offer insight into the causes of psychiatric disorders, general wellbeing10, and longevity. The broad trait of neuroticism is composed of narrower traits, or factors. It was recently discovered that, whereas higher scores on the broad trait of neuroticism are associated with earlier death, higher scores on a worry/vulnerability factor are associated with living longer11. Here, we examine the genetic architectures of two neuroticism factors--worry/vulnerability and anxiety/tension--and how they contrast with the architecture of the general factor of neuroticism. We show that, whereas the polygenic load for general factor of neuroticism is associated with an increased risk of coronary artery disease (CAD), major depressive disorder, and poorer self-rated health, the genetic variants associated with high levels of the anxiety/tension and worry/vulnerability factors are associated with affluence, higher cognitive ability, better self-rated health, and longer life. We also identify the first genes associated with factors of neuroticism that are linked with these positive outcomes that show no relationship with the general factor of neuroticism.

genetics

Genetics and epigenetic alterations of hexaploid early generation derived from hybrid between Brassica napus and B. oleracea

Good fertility was observed previously in hexaploid derived from hybrid (ACC) between calona Zhongshuang 9(Brassica napus, 2n = 38, AACC) and kale SWU01 (B. oleracea var. acephala, 2n = 18, CC). However, the mechanism to underlying the character is unknown. In the present study, genetic and epigenetic alterations of S0, 6 S1, and 18 of their S2 progenies with hexaploid chromosome conformation (20A + 36C) were selected to compare with ACC and their parental species. 13.08% and 26.45% polymorphism alleles different from two parental species were identified in ACC via 58 SSR (simple sequence repeats) and 14 MSAP (methylation sensitive amplified polymorphism), respectively. 33.74% new alleles in DNA methylation, but not in DNA sequence were detected in S0 after chromosome doubling of ACC. DNA profilling revealed a little genetic but much epigenetic differences among S0, S1 and S2 generations. Genetic alteration was relatively stable, because only 8.09% and 3.21% alleles inheriated from ACC were changed in S2 and S1, respectively. While on average of 52.44 {+/-} 5.32% DNA methylation site inherited from ACC were detected in S1, and 41.52 {+/-} 9.04% in S2 due to dramatic epigenetic variance among early generations. New DNA methylation sites occurred in S0 would inheritated into the successive generations, but the frequency was decreased because some new site might be recovered. It demonstrated that much DNA methylation but a little DNA sequence variance was occurred in hexaploid early generation.

genetics

Mapping in vivo genetic interactomics through Cpf1 crRNA array screening

Genetic interactions lay the foundation of biological networks in virtually all organisms. Due to the complexity of mammalian genomes and cellular architectures, unbiased mapping of genetic interactions in vivo is challenging. Cpf1 is a single effector RNA-guided nuclease that enables multiplexed genome editing using crRNA arrays. Here we designed a Cpf1 crRNA array library targeting all pairwise permutations of the most significantly mutated nononcogenes, and performed double knockout screens in mice using a model of malignant transformation as well as a model of metastasis. CrRNA array sequencing revealed a quantitative landscape of all single and double knockouts. Enrichment, synergy and clonal analyses identified many unpredicted drivers and co-drivers of transformation and metastasis, with epigenetic factors as hubs of these highly connected networks. Our study demonstrates a powerful yet simple approach for in vivo mapping of unbiased genetic interactomes in mammalian species at a phenotypic level.

genetics

Complex genetic patterns in human arise from a simple range-expansion model over continental landmasses

Although it is generally accepted that geography is a major factor shaping human genetic differentiation, it is still disputed how much of this differentiation is a result of a simple process of isolation-by-distance, and if there are factors generating distinct clusters of genetic similarity. We address this question using a geographically explicit simulation framework coupled with an Approximate Bayesian Computation approach. Based on six simple summary statistics only, we estimated the most probable demographic parameters that shaped modern human evolution under an isolation by distance scenario, and found these were the following: an initial population in East Africa spread and grew from 4000 individuals to 5.7 million in about 132 000 years. Subsequent simulations with these estimates followed by cluster analyses produced results nearly identical to those obtained in real data. Thus, a simple diffusion model from East Africa explains a large portion of the genetic diversity patterns observed in modern humans. We argue that a model of isolation by distance along the continental landmasses might be the relevant null model to use when investigating selective effects in humans and probably many other species.

genetics

The maternal genetic history of the Angolan Namib Desert: a key region for understanding the peopling of southern Africa

Southern Angola is a poorly studied region, inhabited by populations that have been associated with different migratory movements into southern Africa. Besides the long-standing presence of indigenous Kxa-speaking foragers and the more recent arrival of Bantu-speaking pastoralists, ethnographic and linguistic studies have suggested that other pre-Bantu communities were also present in the Namib desert, including peripatetic groups like the Kwepe (formerly Kwadi speakers), Twa and Kwisi. Here we evaluate previous peopling hypotheses by analyzing the relationships between seven groups from the Namib desert (Kuvale, Himba, Tjimba, Kwisi, Twa, Kwepe) and Kunene Province (!Xun), based on newly collected linguistic data and 295 complete mtDNA genomes. We found that: i) all groups from the Namib desert have genealogically-consistent matriclanic systems that had a strong impact on their maternal genetic structure by enhancing genetic drift and population differentiation; ii) the dominant pastoral groups represented by the Kuvale and Himba were part of a Bantu proto-population that also included the ancestors of present-day Damara and Herero peoples from Namibia; iii) Tjimba are closely related to the Himba; iv) the Kwepe, Twa and Kwisi have a divergent Bantu-related mtDNA profile and probably stem from a single population that does not show clear signs of being a pre-Bantu indigenous group. Taken together, our results suggest that the maternal genetic structure of the different groups from the Namib desert is largely derived from endogamous Bantu peoples, and that their social stratification and different subsistence patterns are not indicative of remnant groups, but reflect Bantu-internal variation and ethnogenesis.

genetics

A comprehensive map of genetic variation in the world’s largest ethnic group - Han Chinese

As are most non-European populations around the globe, the Han Chinese are relatively understudied in population and medical genetics studies. From low-coverage whole-genome sequencing of 11,670 Han Chinese women we present a catalog of 25,057,223 variants, including 548,401 novel variants that are seen at least 10 times in our dataset. Individuals from our study come from 19 out of 22 provinces across China, allowing us to study population structure, genetic ancestry, and local adaptation in Han Chinese. We identify previously unrecognized population structure along the East-West axis of China and report unique signals of admixture across geographical space, such as European influences among the Northwestern provinces of China. Finally, we identified a number of highly differentiated loci, indicative of local adaptation in the Han Chinese. In particular, we detected extreme differentiation among the Han Chinese at MTHFR, ADH7, and FADS loci, suggesting that these loci may not be specifically selected in Tibetan and Inuit populations as previously suggested. On the other hand, we find that Neandertal ancestry does not vary significantly across the provinces, consistent with admixture prior to the dispersal of modern Han Chinese. Furthermore, contrary to a previous report, Neandertal ancestry does not explain a significant amount of heritability in depression. Our findings provide the largest genetic data set so far made available for Han Chinese and provide insights into the history and population structure of the worlds largest ethnic group.

genetics

Genetics and genomics of social behaviour in a chicken model

The identification of genes affecting behaviour can be problematic, yet their identification allows a raft of possibilities. Sociality and social behaviour can have multiple definitions, though at its core it is the desire to seek contact with con- or hetero-specifics. The identification of genes affecting sociality can therefore give insights into the maintenance and establishment of sociality. In this study we used the combination of an advanced intercross between wild and domestic chickens with a combined QTL and eQTL genetical genomics approach to identify genes for social reinstatement (SR) behaviour. A total of 24 SR QTL were identified and overlaid with over 600 eQTL obtained from the same birds using hypothalamus tissue. Correlations between overlapping QTL and eQTL indicated 5 strong candidate genes, with the gene TTRAP being strongly significantly correlated with multiple aspects of SR behaviour, as well as possessing a highly significant eQTL. The distribution of eQTL can also indicate the genetic mechanisms underlying domestication itself. Multiple eQTL were found to in discrete clusters, however tests for pleiotropy show that these blocks were primarily linked in origin. This suggests that clustered genetic modules, rather than pure pleiotropy (as hypothesised by the neural crest theory) appears to be driving domestication in the chicken.

genetics

A family-based method for leveraging random genetic variation to identify variance-controlling loci

The propensity of a trait to vary within a population may have evolutionary, ecological, or clinical significance. In the present study we deploy sibling models to offer a novel and unbiased way to ascertain loci associated with the extent to which phenotypes vary (variance-controlling quantitative trait loci, or vQTLs). Previous methods for vQTL-mapping either exclude genetically related individuals or treat genetic relatedness among individuals as a complicating factor addressed by adjusting estimates for non-independence in phenotypes. The present method uses genetic relatedness as a tool to obtain unbiased estimates of variance effects rather than as a nuisance. The family-based approach, which utilizes random variation between siblings in minor allele counts at a locus, also allows controls for parental genotype, mean effects, and non-linear (dominance) effects that may spuriously appear to generate variation.\n\nSimulations show that the approach performs equally well as two existing methods (squared Z-score and DGLM) in controlling type I error rates when there is no unobserved confounding, and performs significantly better than these methods in the presence of confounding. Using height and BMI as empirical applications, we investigate SNPs that alter within-family variation in height and BMI, as well as pathways that appear to be enriched. One significant SNP for BMI variability, in the MAST4 gene, replicated. Pathway analysis revealed one gene set, encoding members of several signaling pathways related to gap junction function, which appears significantly enriched for associations with within-family height variation in both datasets (while not enriched in analysis of mean levels). We recommend approximating laboratory random assignment of genotype using family data and more careful attention to the possible conflation of mean and variance effects.

genetics

Genome-wide analysis of risk-taking behaviour and cross-disorder genetic correlations in 116,255 individuals from the UK Biobank cohort

Risk-taking behaviour is a key component of several psychiatric disorders and could influence lifestyle choices such as smoking, alcohol use and diet. As a phenotype, risk-taking behaviour therefore fits within a Research Domain Criteria (RDoC) approach, whereby identifying genetic determinants of this trait has the potential to improve our understanding across different psychiatric disorders. Here we report a genome wide association study in 116 255 UK Biobank participants who responded yes/no to the question \"Would you consider yourself a risk-taker?\" Risk-takers (compared to controls) were more likely to be men, smokers and have a history of psychiatric disorder. Genetic loci associated with risk-taking behaviour were identified on chromosomes 3 (rs13084531) and 6 (rs9379971). The effects of both lead SNPs were comparable between men and women. The chromosome 3 locus highlights CADM2, previously implicated in cognitive and executive functions, but the chromosome 6 locus is challenging to interpret due to the complexity of the HLA region. Risk-taking behaviour shared significant genetic risk with schizophrenia, bipolar disorder, attention deficit hyperactivity disorder and post-traumatic stress disorder, as well as with smoking and total obesity. Despite being based on only a single question, this study furthers our understanding of the biology of risk-taking behaviour, a trait which has a major impact on a range of common physical and mental health disorders.

genetics