Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Spider web DNA: a new spin on noninvasive genetics of predator and prey

Noninvasive genetic approaches enable biomonitoring without the need to directly observe or disturb target organisms. Environmental DNA (eDNA) methods have recently extended this approach by assaying genetic material within bulk environmental samples without a priori knowledge about the presence of target biological material. This paper describes a novel and promising source of noninvasive spider DNA and insect eDNA from spider webs. Using black widow spiders (Latrodectus spp.) fed with house crickets (Acheta domesticus), we successfully extracted and amplified mitochondrial DNA sequences of both spider and prey from spider web. Detectability of spider DNA did not differ between assays with amplicon sizes from 135 to 497 base pairs. Spider DNA and prey eDNA remained detectable at least 88 days after living organisms were no longer present on the web. Spider web DNA may be an important tool in conservation research, pest management, biogeography studies, and biodiversity assessments.

Genetics

Testing for genetic associations in arbitrarily structured populations

We present a new statistical test of association between a trait (either quantitative or binary) and genetic markers, which we theoretically and practically prove to be robust to arbitrarily complex population structure. The statistical test involves a set of parameters that can be directly estimated from large-scale genotyping data, such as that measured in genome-wide associations studies (GWAS). We also derive a new set of methodologies, called a genotype-conditional association test (GCAT), shown to provide accurate association tests in populations with complex structures, manifested in both the genetic and environmental contributions to the trait. We demonstrate the proposed method on a large simulation study and on the Northern Finland Birth Cohort study. In the Finland study, we identify several new significant loci that other methods do not detect. Our proposed framework provides a substantially different approach to the problem from existing methods. We provide some discussion on its similarities and differences with the linear mixed model and principal component approaches.

Genetics

Genetic Variation, Not Cell Type of Origin, Underlies Regulatory Differences in iPSCs

The advent of induced pluripotent stem cells (iPSCs)1 revolutionized Human Genetics by allowing us to generate pluripotent cells from easily accessible somatic tissues. This technology can have immense implications for regenerative medicine, but iPSCs also represent a paradigm shift in the study of complex human phenotypes, including gene regulation and disease2-5. Yet, an unresolved caveat of the iPSC model system is the extent to which reprogrammed iPSCs retain residual phenotypes from their precursor somatic cells. To directly address this issue, we used an effective study design to compare regulatory phenotypes between iPSCs derived from two types of commonly used somatic precursor cells. We show that the cell type of origin only minimally affects gene expression levels and DNA methylation in iPSCs. Instead, genetic variation is the main driver of regulatory differences between iPSCs of different donors.

Genetics

An accurate genetic clock

Molecular clocks give \"Time to most recent common ancestor\" TMRCA of genetic trees. By Watson-Galton17 most lineages terminate, with a few overrepresented singular lineages generated by W. Hamiltons \"kin selection\"13. Applying current methods to this non-uniform branching produces greatly exaggerated TMRCA. We introduce an inhomogenous stochastic process which detects singular lineages by asymmetries, whose reduction gives true TMRCA. This implies a new method for computing mutation rates. Despite low rates similar to mitosis data, reduction implies younger TMRCA, with smaller errors. We establish accuracy by a comparison across a wide range of time, indeed this is only clock giving consistent results for both short and long term times. In particular we show that the dominant European y-haplotypes R1a1a & R1b1a2, expand from c3700BC, not reaching Anatolia before c3300BC. While this contradicts current clocks which date R1b1a2 to either the Neolithic Near East4 or Paleo-Europe20, our dates support recent genetic analysis of ancient skeletons by Reich23.

Genetics

A Perspective on Interaction Tests in Genetic Association Studies

The identification of gene-gene and gene-environment interaction in human traits and diseases is an active area of research that generates high expectation, and most often lead to high disappointment. This is partly explained by a misunderstanding of some of the inherent characteristics of interaction effects. Here, I untangle several theoretical aspects of standard regression-based interaction tests in genetic association studies. In particular, I discuss variables coding scheme, interpretation of effect estimate, power, and estimation of variance explained in regard of various hypothetical interaction patterns. I show first that the simplest biological interaction models--in which the magnitude of a genetic effect depends on a common exposure--are among the most difficult to identify. Then, I demonstrate the demerits of the current strategy to evaluate the contribution of interaction effects to the variance of quantitative outcomes and argue for the use of new approaches to overcome these issues. Finally I explore the advantages and limitations of multivariate models when testing for interaction between multiple SNPs and/or multiple exposures, using either a joint test, or a test of interaction based on risk score. Theoretical and simulated examples presented along the manuscript demonstrate that the application of these methods can provide a new perspective on the role of interaction in multifactorial traits.

Genetics

The impact of host metapopulation structure on the population genetics of colonizing bacteria

Many key bacterial pathogens are frequently carried asymptomatically, and the emergence and spread of these opportunistic pathogens can be driven, or mitigated, via demographic changes within the host population. These inter-host transmission dynamics combine with basic evolutionary parameters such as rates of mutation and recombination, population size and selection, to shape the genetic diversity within bacterial populations. Whilst many studies have focused on how molecular processes underpin bacterial population structure, the impact of host migration and the connectivity of the local populations has received far less attention. A stochastic neutral model incorporating heightened local transmission has been previously shown to fit closely with genetic data for several bacterial species. However, this model did not incorporate transmission limiting population stratification, nor the possibility of migration of strains between subpopulations, which we address here by presenting an extended model. The model captures the observed population patterns for the common nosocomial pathogens Staphylococcus epidermidis and Enterococcus faecalis, while Staphylococcus aureus and Enterococcus faecium display deviations attributable to adaptation. It is demonstrated analytically and numerically that expected strain relatedness may either increase or decrease as a function of increasing migration rate between subpopulations, being a complex function of the rate at which microepidemics occur in the metapopulation. Moreover, it is shown that in a structured population markedly different rates of evolution may lead to indistinguishable patterns of relatedness among bacterial strains; caution is thus required when drawing evolution inference in these cases.

Genetics

Hypothesis-free identification of modulators of genetic risk factors

Genetic risk factors often localize in non-coding regions of the genome with unknown effects on disease etiology. Expression quantitative trait loci (eQTLs) help to explain the regulatory mechanisms underlying the association of genetic risk factors with disease. More mechanistic insights can be derived from knowledge of the context, such as cell type or the activity of signaling pathways, influencing the nature and strength of eQTLs. Here, we generated peripheral blood RNA-seq data from 2,116 unrelated Dutch individuals and systematically identified these context-dependent eQTLs using a hypothesis-free strategy that does not require prior knowledge on the identity of the modifiers. Out of the 23,060 significant cis-regulated genes (false discovery rate < 0.05), 2,743 genes (12%) show context-dependent eQTL effects. The majority of those were influenced by cell type composition, revealing eQTLs that are particularly strong in cell types such as CD4+ T-cells, erythrocytes, and even lowly abundant eosinophils. A set of 145 cis-eQTLs were influenced by the activity of the type I interferon signaling pathway and we identified several cis-eQTLs that are modulated by specific transcription factors that bind to the eQTL SNPs. This demonstrates that large-scale eQTL studies in unchallenged individuals can complement perturbation experiments to gain better insight in regulatory networks and their stimuli.

Genetics

Simple penalties on maximum likelihood estimates of genetic parameters to reduce sampling variation

Multivariate estimates of genetic parameters are subject to substantial sampling variation, especially for smaller data sets and more than a few traits. A simple modification of standard, maximum likelihood procedures for multivariate analyses to estimate genetic covariances is described, which can improve estimates by substantially reducing their sampling variances. This is achieved maximizing the likelihood subject to a penalty. Borrowing from Bayesian principles, we propose a mild, default penalty - derived assuming a Beta distribution of scale-free functions of the covariance components to be estimated - rather than laboriously attempting to determine the stringency of penalization from the data. An extensive simulation study is presented demonstrating that such penalties can yield very worthwhile reductions in loss, i.e. the difference from population values, for a wide range of scenarios and without distorting estimates of phenotypic covariances. Moreover, mild default penalties tend not to increase loss in difficult cases and, on average, achieve reductions in loss of similar magnitude than computationally demanding schemes to optimize the degree of penalization. Pertinent details required for the adaptation of standard algorithms to locate the maximum of the likelihood function are outlined.

Genetics

Translational plasticity facilitates the accumulation of nonsense genetic variants in the human population

Genetic variants that disrupt protein-coding DNA are ubiquitous in the human population, with ~100 such loss-of-function variants per individual. While most loss-of-function variants are rare, a subset have risen to high frequency and occur in a homozygous state in healthy individuals. It is unknown why these common variants are well-tolerated, even though some affect essential genes implicated in Mendelian disease. Here, we combine genomic, proteomic, and biochemical data to demonstrate that many common nonsense variants do not ablate protein production from their host genes. We provide computational and experimental evidence for diverse mechanisms of gene rescue, including alternative splicing, stop codon readthrough, alternative translation initiation, and C-terminal truncation. Our results suggest a molecular explanation for the mild fitness costs of many common nonsense variants, and indicate that translational plasticity plays a prominent role in shaping human genetic diversity.

Genetics

Genetic contributions to self-reported tiredness

Self-reported tiredness and low energy, often called fatigue, is associated with poorer physical and mental health. Twin studies have indicated that this has a heritability between 6% and 50%. In the UK Biobank sample (N = 108 976) we carried out a genome-wide association study of responses to the question, \"Over the last two weeks, how often have you felt tired or had little energy?\" Univariate GCTA-GREML found that the proportion of variance explained by all common SNPs for this tiredness question was 8.4% (SE = 0.6%). GWAS identified one genome-wide significant hit (Affymetrix id 1:64178756_C_T; p = 1.36 x 10-11). LD score regression and polygenic profile analysis were used to test for pleiotropy between tiredness and up to 28 physical and mental health traits from GWAS consortia. Significant genetic correlations were identified between tiredness and BMI, HDL cholesterol, forced expiratory volume, grip strength, HbA1c, longevity, obesity, self-rated health, smoking status, triglycerides, type 2 diabetes, waist-hip ratio, ADHD, bipolar disorder, major depressive disorder, neuroticism, schizophrenia, and verbal-numerical reasoning (absolute rg effect sizes between 0.11 and 0.78). Significant associations were identified between tiredness phenotypic scores and polygenic profile scores for BMI, HDL cholesterol, LDL cholesterol, coronary artery disease, HbA1c, height, obesity, smoking status, triglycerides, type 2 diabetes, and waist-hip ratio, childhood cognitive ability, neuroticism, bipolar disorder, major depressive disorder, and schizophrenia (standardised {beta}s between -0.016 and 0.03). These results suggest that tiredness is a partly-heritable, heterogeneous and complex phenomenon that is phenotypically and genetically associated with affective, cognitive, personality, and physiological processes.\n\n\"Hech, sirs! But Im wabbit, Im back frae the toon;\n\nI haena dune pechin--jist let me sit doon.\n\nFrom Glesca\n\nBy William Dixon Cocker (1882-1970)

Genetics

Using Y chromosomal haplogroups in genetic association studies and suggested implications

Y chromosomal (Y-DNA) haplogroups are more widely used in population genetics than in genetic epidemiology, although associations between Y-DNA haplogroups and several traits (including cardio-metabolic traits) have been reported. In apparently homogeneous populations, there is still Y-DNA haplogroup variation which will result from population history. Therefore, hidden stratification and/or differential phenotypic effects by Y-DNA haplogroups could exist. To test this, we hypothesised that stratifying individuals according to their Y-DNA haplogroups before testing associations between autosomal SNPs and phenotypes will yield difference in association. For proof of concept, we derived Y-DNA haplogroups from 6,537 males from two epidemiological cohorts, ALSPAC (N=5,080, 816 Y-DNA SNPs) and 1958 Birth Cohort (N=1,457, 1,849 Y-DNA SNPs). For illustration, we studied well-known associations between 32 SNPs and body mass index (BMI), including associations involving FTO SNPs. Overall, no association was replicated in both cohorts when Y-DNA haplogroups were considered and this suggests that, for BMI at least, there is little evidence of differences in phenotype or gene association by Y-DNA structure. Further studies using other traits, Phenome-wide association studies (PheWAS), haplogroups and/or autosomal SNPs are required to test the generalisability of this approach.

Genetics

The genetic structure of the world’s first farmers

We report genome-wide ancient DNA from 44 ancient Near Easterners ranging in time between ~12,000-1,400 BCE, from Natufian hunter-gatherers to Bronze Age farmers. We show that the earliest populations of the Near East derived around half their ancestry from a Basal Eurasian lineage that had little if any Neanderthal admixture and that separated from other non-African lineages prior to their separation from each other. The first farmers of the southern Levant (Israel and Jordan) and Zagros Mountains (Iran) were strongly genetically differentiated, and each descended from local hunter-gatherers. By the time of the Bronze Age, these two populations and Anatolian-related farmers had mixed with each other and with the hunter-gatherers of Europe to drastically reduce genetic differentiation. The impact of the Near Eastern farmers extended beyond the Near East: farmers related to those of Anatolia spread westward into Europe; farmers related to those of the Levant spread southward into East Africa; farmers related to those from Iran spread northward into the Eurasian steppe; and people related to both the early farmers of Iran and to the pastoralists of the Eurasian steppe spread eastward into South Asia.

Genetics

Using imputed genotype data in the joint score tests for genetic association and gene-environment interactions in case-control studies

BackgroundGenome-wide association studies (GWAS) are now routinely imputed for untyped SNPs based on various powerful statistical algorithms for imputation trained on reference datasets. The use of predicted allele count for imputed SNPs as the dosage variable is known to produce valid score test for genetic association.\n\nMethodsIn this paper, we investigate how to best handle imputed SNPs in various modern complex tests for genetic association incorporating gene-environment interactions. We focus on case-control association studies where inference in an underlying logistic regression model can be performed using alternative methods that rely on varying degree on an assumption of gene-environment independence in the underlying population. As increasingly large scale GWAS are being performed through consortia effort where it is preferable to share only summary-level information across studies, we also describe simple mechanisms for implementing score-tests based on standard meta-analysis of \"one-step\" maximum-likelihood estimates across studies.\n\nResultsApplications of the methods in simulation studies and a dataset from genome-wide association study of lung cancer illustrate ability of the proposed methods to maintain type-I error rates for underlying testing procedures. For analysis of imputed SNPs, similar to typed SNPs, retrospective methods can lead to considerable efficiency gain for modeling of gene-environment interactions under the assumption of gene-environment independence.\n\nConclusionsProposed methods allow valid analysis of imputed SNPs in case-control studies of gene-environment interaction using alternative strategies that had been earlier available only for genotyped SNPs.

Genetics

High-resolution DNA accessibility profiles increase the discovery and interpretability of genetic associations

Genetic risk for common autoimmune diseases is influenced by hundreds of small effect, mostly non-coding variants, enriched in regulatory regions active in adaptive-immune cell types. DNaseI hypersensitivity sites (DHSs) are a genomic mark for regulatory DNA. Here, we generated a single DHSs annotation from fifteen deeply sequenced DNase-seq experiments in adaptive-immune as well as non-immune cell types. Using this annotation we quantified accessibility across cell types in a matrix format amenable to statistical analysis, deduced the subset of DHSs unique to adaptive-immune cell types, and grouped DHSs by cell-type accessibility profiles. Measuring enrichment with cell-type-specific TF binding sites as well as proximal gene expression and function, we show that accessibility profiles grouped DHSs into coherent regulatory functions. Using the adaptive-immune-specific DHSs as input (0.37% of genome), we associated DHSs to six autoimmune diseases with GWAS data. Associated loci showed higher replication rates when compared to loci identified by GWAS or by considering all DHSs, allowing the additional discovery of 327 loci (FDR<0.005) below typical GWAS significance threshold, 52 of which are novel and replicating discoveries. Finally, we integrated DHS associations from six autoimmune diseases, using a network model (bird-eye view) and a regulatory Manhattan plot schema (per locus). Taken together, we described and validated a strategy to leverage finely resolved regulatory priors, enhancing the discovery, interpretability, and resolution of genetic associations, and providing actionable insights for follow up work.

Genetics

Chromatin Landscapes and Genetic Risk For Juvenile Idiopathic Arthritis

Juvenile idiopathic arthritis (JIA) is considered to be an autoimmune disease mediated by interactions between genes and the environment. To gain a better understanding of the cellular basis for genetic risk, we studied known JIA genetic risk loci, the majority of which are located in non-coding regions, in human neutrophils and CD4 primary T cells to identify genes and functional elements located within those risk loci. We analyzed RNA-Seq data, H3K27ac and H3K4me1 chromatin immunoprecipitation-sequencing (ChIP-Seq) data, and previously published chromatin interaction analysis by paired-end tag sequencing (ChIA-PET) data to characterize the chromatin landscapes within the know JIA-associated risk loci. In both neutrophils and primary CD4+ T cells, the majority of the JIA-associated LD blocks contained H3K27ac and/or H3K4me1 marks. These LD blocks were also binding sites for a small group of transcription factors, particularly in neutrophils. Furthermore, these regions showed abundant intronic and intergenic transcription in neutrophils. In neutrophils, none of the genes that were differentially expressed between untreated JIA patients and healthy children was located within the JIA risk LD blocks. In CD4+ T cells, multiple genes, including HLA-DQA1, HLA-DQB2, TRAF1, and IRF1 were associated with the long-distance interacting regions within the LD regions as determined from ChIA-PET data. These findings suggest that aberrant transcriptional control is the underlying pathogenic mechanism in JIA. Furthermore, these findings demonstrate the challenges of identifying the actual causal variants within complex genomic/chromatin landscapes.

Genetics

Educational attainment and personality are genetically intertwined

Heritable variance in psychological traits may reflect genetic and biological processes that are not necessarily specific to these particular traits but pertain to a broader range of phenotypes. We tested the possibility that Five-Factor Model personality domains and their 30 facets, as rated by people themselves and their knowledgeable informants, reflect polygenic influences that have been previously associated with educational attainment. In a sample of over 3,000 adult Estonians, polygenic scores for educational attainment (EPS; interpretable as estimates of molecular genetic propensity for education) were correlated with various personality traits, particularly from the Neuroticism and Openness domains. The correlations of personality traits with phenotypic educational attainment closely mirrored their correlations with EPS. Moreover, EPS predicted an aggregate personality trait tailored to capture maximum amount of variance in educational attainment almost as strongly as it predicted the attainment itself. We discuss possible interpretations and implications of these findings.

Genetics

biMM: Efficient estimation of genetic variances andcovariances for cohorts with high-dimensional phenotype measurements

Genetic research utilizes a decomposition of trait variances and covariances into genetic and environmental parts. Our software package biMM is a computationally efficient implementation of a bivariate linear mixed model for settings where hundreds of traits have been measured on partially overlapping sets of individuals.\n\nAvailabilityImplementation in R freely available at www.iki.fi/mpirinen.

genetics

Genetic variation and gene expression across multiple tissues and developmental stages in a non-human primate

By analyzing multi-tissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalogue of expression quantitative trait loci (eQTLs) in a non-human primate model. This catalogue contains more genome-wide significant eQTLs, per sample, than comparable human resources, and reveals sex and age-related expression patterns. Findings include a master regulatory locus that likely plays a role in immune function, and a locus regulating hippocampal long non-coding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.

genetics