Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Subset-Based Analysis using Gene-Environment Interactions for Discovery of Genetic Associations across Multiple Studies or Phenotypes

ObjectivesClassical methods for combining summary data from genome-wide association studies (GWAS) only use marginal genetic effects and power can be compromised in the presence of heterogeneity. We aim to enhance the discovery of novel associated loci in the presence of heterogeneity of genetic effects in sub-groups defined by an environmental factor.\n\nMethodsWe present a p-value Assisted Subset Testing for Associations (pASTA) framework that generalizes the previously proposed association analysis based on subsets (ASSET) method by incorporating gene-environment (G-E) interactions into the testing procedure. We conduct simulation studies and provide two data examples.\n\nResultsSimulation studies show that our proposal is more powerful than methods based on marginal associations in the presence of G-E interactions and maintains comparable power even in their absence. Both data examples demonstrate that our method can increase power to detect overall genetic associations and identify novel studies/phenotypes that contribute to the association.\n\nConclusionsOur proposed method can be a useful screening tool to identify candidate single nucleotide polymorphisms (SNPs) that are potentially associated with the trait(s) of interest for further validation. It also allows researchers to determine the most probable subset of traits that exhibit genetic associations in addition to the enhancement of power.

genetics

Phenome-wide association analysis of LDL-cholesterol lowering genetic variants in PCSK9

BackgroundWe characterised the phenotypic consequence of genetic variation at the PCSK9 locus and compared findings with recent trials of pharmacological inhibitors of PCSK9.\n\nMethodsPublished and individual participant level data (300,000+ participants) were combined to construct a weighted PCSK9 gene-centric score (GS). Fourteen randomized placebo controlled PCSK9 inhibitor trials were included, providing data on 79,578 participants. Results were scaled to a one mmol/L lower LDL-C concentration\n\nResultsThe PCSK9 GS (comprising 4 SNPs) associations with plasma lipid and apolipoprotein levels were consistent in direction with treatment effects. The GS odds ratio (OR) for myocardial infarction (MI) was 0.53 (95%CI 0.42; 0.68), compared to a PCSK9 inhibitor effect of 0.90 (95%CI 0.86; 0.93). For ischemic stroke ORs were 0.84 (95%CI 0.57; 1.22) for the GS, compared to 0.85 (95%CI 0.78; 0.93) in the drug trials. ORs with type 2 diabetes mellitus (T2DM) were 1.29 (95% CI 1.11; 1.50) for the GS, as compared to 1.00 (95%CI 0.96; 1.04) for incident T2DM in PCSK9 inhibitor trials. No genetic associations were observed for cancer, heart failure, atrial fibrillation, chronic obstructive pulmonary disease, or Alzheimers disease - outcomes for which large-scale trial data were unavailable.\n\nConclusionsGenetic variation at the PCSK9 locus recapitulates the effects of therapeutic inhibition of PCSK9 on major blood lipid fractions and MI. Apparent discordance between genetic associations and trial outcome for T2DM might be explained lack by a of statistical precision, or differences in the nature and duration of genetic versus pharmacological perturbation of PCSK9.\n\nFundingThis research was funded by the British Heart Foundation (SP/13/6/30554, RG/10/12/28456, FS/18/23/33512), UCL Hospitals NIHR Biomedical Research Centre, by the Rosetrees and Stoneygate Trusts.\n\nCondensed abstractEvidence on the long-term efficacy and safety of therapeutic inhibition of PCSK9 is lacking. To explore potential long-term effects of PCSK9 inhibition, we characterised the phenotypic consequence of LDL-cholesterol lowering variants at the PCSK9 locus. A PCSK9 gene score comprising 4 SNPs recapitulated the effects of therapeutic inhibition of PCSK9 on major blood lipid fractions and risk of myocardial infarction, and was associated with an increased risk of type 2 diabetes. No associations with safety outcomes such as cancer, COPD, Alzheimers disease or atrial fibrillation were identified. Our findings suggest PCSK9 inhibition may be safe and effective during prolonged use.

genetics

Genetic Modifiers of Pathogenic LRRK2 G2019S Neurodegeneration in Drosophila

Disease phenotypes can be highly variable among individuals with the same pathogenic mutation. There is increasing evidence that background genetic variation is a strong driver of disease variability in addition to the influence of environment. To understand the genotype-phenotype relationship that determines the expressivity of a pathogenic mutation, a large number of backgrounds must be studied. This can be efficiently achieved using model organism collections such as the Drosophila Genetic Reference Panel (DGRP). Here, we used the DGRP to assess the variability of locomotor dysfunction in a LRRK2 G2019S Drosophila melanogaster model of Parkinsons disease. We find substantial variability in the LRRK2 G2019S locomotor phenotype in different DGRP backgrounds. A genome-wide association study for candidate genetic modifiers reveals 177 genes that drive wide phenotypic variation, including 19 top association genes. Genes involved in the outgrowth and regulation of neuronal projections are enriched in these candidate modifiers. RNAi functional testing of the top association and neuronal projection-related genes reveals that pros, pbl, ct and CG33506 significantly modify age-related dopamine neuron loss and associated locomotor dysfunction in the Drosophila LRRK2 G2019S model. These results demonstrate how natural genetic variation can be used as a powerful tool to identify genes that modify disease-related phenotypes. We report novel candidate modifier genes for LRRK2 G2019S that may be used to interrogate the link between LRRK2, neurite regulation and neuronal degeneration in Parkinsons disease.

genetics

Analysis of genetic control and QTL mapping of essential wheat grain quality traits in a recombinant inbred population

Wheat cultivars are genetically crossed for improving end use quality for apt traits as per need of baking industry and broad consumers preferences. The processing and baking qualities of bread wheat underlie into genetic make-up of a variety and influence by environmental factors and their interactions. WL711 and C306 derived recombinant inbred lines (RILs) population of 206 was used for phenotyping of quality related traits in three different environmental conditions. The genetic analysis of quality traits showed considerable variation for measurable quality traits with normal distribution and transgressive segregation across the years. From the 206 RIL, few RILs found to be superior to those of the parental cultivars for key quality traitsindicating their potential usefor improvement of end use quality and also suggestingprobability of finding new alleles and allelic combinations from the RIL population. A genetic linkage map including 346 markers was constructed withtotal map distance of 4526.8cM andinterval distance between adjacent markersof 12.9cM. Mapping analysis identified 38 putative QTLs for 13 quality related traits with QTLs explaining 7.9% - 16.8% phenotypic variation spanning over 14 chromosomes i.e. 1A, 1B, 1D, 2A, 2D, 3B, 3D, 4A, 4B, 4D, 5D, 6A, 7A and 7B. Major novel QTLs regions for quality traits have been identified on several chromosome in studied RIL population posing their potential role in marker assisted selection for better bread making quality after validation.

genetics

Longevity defined as top 10% survivors is transmitted as a quantitative genetic trait: results from large three-generation datasets

Survival to extreme ages clusters within families. However, identifying genetic loci conferring longevity and low morbidity in such longevous families is challenging. There is debate concerning the survival percentile that best isolates the genetic component in longevity. Here, we use three-generational mortality data from two large datasets, UPDB (US) and LINKS (Netherlands). We studied 21,046 unselected families containing index persons, their parents, siblings, spouses, and children, comprising 321,687 individuals. Our analyses provide strong evidence that longevity is transmitted as a quantitative genetic trait among survivors up to the top 10% of their birth cohort. We subsequently showed a survival advantage, mounting to 31%, for individuals with top 10% surviving first and second-degree relatives in both databases and across generations, even in the presence of non-longevous parents. To guide future genetic studies, we suggest to base case selection on top 10% survivors of their birth cohort with equally long-lived family members.

genetics

Systematic classification of shared components of genetic risk for common human diseases

Disease classification is fundamental to clinical practice, but current taxonomies do not necessarily reflect the pathophysiological processes that are common or unique to different disorders, such as those determined by genetic risk factors. Here, we use routine healthcare data from the 500,000 participants in the UK Biobank to map genome-wide associations across 19,628 diagnostic terms. We find that 3,510 independent genetic risk loci affect multiple clinical phenotypes, which we cluster into 629 distinct disease association profiles. We use multiple approaches to link clusters to different underlying biological pathways and show how these clusters define the genetic architecture of common medical conditions, including hypertension and immune-mediated diseases. Finally, we demonstrate how clusters can be utilised to re-define disease relationships and to inform therapeutic strategies.\n\nOne sentence summarySystematic classification of genetic risk factors reveals molecular connectivity of human diseases with clinical implications

genetics

Genome Wide Meta-Analysis identifies new loci associated with cardiac phenotypes and uncovers a common genetic signature shared by heart function and Alzheimer’s disease

AimsEchocardiography has become an indispensable tool for the study of heart performance, improving the monitoring of individuals with cardiac diseases. Diverse genetic factors associated with echocardiographic measures of heart structure and functions have been previously reported. The impact of several apoptotic genes in heart development identified in experimental models prompted us to assess their potential association with indicators of human cardiac function. This study started with the aim to investigate the possible association of variants of apoptotic genes with echocardiographic traits and to identify new genetic markers associated with cardiac function.\n\nMethods and resultsGenome wide data from different studies were obtained from public repositories. After quality control and imputation, association analyses confirm the role of caspases and other apoptosis related genes with cardiac phenotypes. Moreover, enrichment analysis showed an over-representation of genes, including some apoptotic regulators, associated with Alzheimers disease (AD). We further explored this unexpected observation which was confirmed by genetic correlation analyses.\n\nConclusionsOur findings show the association of apoptotic gene variants with echocardiographic indicators of heart function and reveal a novel potential genetic link between echocardiographic measures in healthy populations and cognitive decline later on in life. These findings may have important implications for preventative strategies combating Alzheimers disease.

genetics

Genome-Wide Control of Population Structure and Relatedness in Genetic Association Studies via Linear Mixed Models with Orthogonally Partitioned Structure

Linear mixed models (LMMs) have become the standard approach for genetic association testing in the presence of sample structure. However, the performance of LMMs has primarily been evaluated in relatively homogeneous populations of European ancestry, despite many of the recent genetic association studies including samples from worldwide populations with diverse ancestries. In this paper, we demonstrate that existing LMM methods can have systematic miscalibration of association test statistics genome-wide in samples with heterogenous ancestry, resulting in both increased type-I error rates and a loss of power. Furthermore, we show that this miscalibration arises due to varying allele frequency differences across the genome among populations. To overcome this problem, we developed LMM-OPS, an LMM approach which orthogonally partitions diverse genetic structure into two components: distant population structure and recent genetic relatedness. In simulation studies with real and simulated genotype data, we demonstrate that LMM-OPS is appropriately calibrated in the presence of ancestry heterogeneity and outperforms existing LMM approaches, including EMMAX, GCTA, and GEMMA. We conduct a GWAS of white blood cell (WBC) count in an admixed sample of 3,551 Hispanic/Latino American women from the Womens Health Initiative SNP Health Association Resource where LMM-OPS detects genome-wide significant associations with corresponding p-values that are one or more orders of magnitude smaller than those from competing LMM methods. We also identify a genome-wide significant association with regulatory variant rs2814778 in the DARC gene on chromosome 1, which generalizes to Hispanic/Latino Americans a previous association with reduced WBC count identified in African Americans.

genetics

Genome-wide association studies in Samoans give insight into the genetic etiology of fasting serum lipid levels.

The current understanding of the genetic architecture of lipids has largely come from genome-wide association studies. To date, few studies have examined the genetic architecture of lipids in Polynesians, and none have in Samoans, whose unique population history, including many population bottlenecks, may provide insight into the biological foundations of variation in lipid levels. Here we performed a genome-wide association study of four fasting serum lipid levels: total cholesterol (TC), high-density lipoprotein (HDL), low-density lipoprotein (LDL), and triglycerides (TG) in a sample of 2,849 Samoans, with validation genotyping for associations in a replication cohort comprising 1,798 Samoans and American Samoans. We identified multiple genome-wide significant associations (P < 5 x 10-8) previously seen in other populations - APOA1 with TG, CETP with HDL, and APOE with TC and LDL - and several suggestive associations (P < 1 x 10-5), including an association of variants downstream of MGAT1 and RAB21 with HDL. However, we observed different association signals for variants near APOE than what has been previously reported in non-Polynesian populations. The association with several known lipid loci combined with the newly-identified associations with variants near MGAT1 and RAB21 suggest that while some of the genetic architecture of lipids is shared between Samoans and other populations, part of the genetic architecture may be Polynesian-specific.

genetics

Genetic And Epigenetic Fine Mapping Of Complex Trait Associated Loci In The Human Liver

Deciphering the impact of genetic variation on gene regulation is fundamental to understanding common, complex human diseases. Although histone modifications are important markers of gene regulatory regions of the genome, any specific histone modification has not been assayed in more than a few individuals in the human liver. As a result, the impacts of genetic variation that direct histone modification states in the liver are poorly understood. Here, we generate the most comprehensive genome-wide dataset of two epigenetic marks, H3K4me3 and H3K27ac, and annotate thousands of putative regulatory elements in the human liver. We integrate these findings with genome-wide gene expression data collected from the same human liver tissues and high-resolution promoter-focused chromatin interaction maps collected from human liver-derived HepG2 cells. We demonstrate widespread functional consequences of natural genetic variation on putative regulatory element activity and gene expression levels. Leveraging these extensive datasets, we fine-map a total of 77 GWAS loci that have been associated with at least one complex phenotype. Our results contribute to the repertoire of genes and regulatory mechanisms governing complex disease development and further the basic understanding of genetic and epigenetic regulation of gene expression in the human liver tissue.

genetics

Variably methylated regions in the newborn epigenome: environmental, genetic and combined influences

BackgroundEpigenetic processes, including DNA methylation (DNAm), are among the mechanisms allowing integration of genetic and environmental factors to shape cellular function. While many studies have investigated either environmental or genetic contributions to DNAm, few have assessed their integrated effects. We examined the relative contributions of prenatal environmental factors and genotype on DNA methylation in neonatal blood at variably methylated regions (VMRs), defined as consecutive CpGs showing the highest variability of DNAm in 4 independent cohorts (PREDO, DCHS, UCI, MoBa, N=2,934).\n\nResultsWe used Akaikes information criterion to test which factors best explained variability of methylation in the cohort-specific VMRs: several prenatal environmental factors (E) including maternal demographic, psychosocial and metabolism related phenotypes, genotypes in cis (G), or their additive (G+E) or interaction (GxE) effects. G+E and GxE models consistently best explained variability in DNAm of VMRs across the cohorts, with G explaining the remaining sites best. VMRs best explained by G, GxE or G+E, as well as their associated functional genetic variants (predicted using deep learning algorithms), were located in distinct genomic regions, with different enrichments for transcription and enhancer marks. Genetic variants of not only G and G+E models, but also of variants in GxE models were significantly enriched in genome wide association studies (GWAS) for complex disorders.\n\nConclusionGenetic and environmental factors in combination best explain DNAm at VMRs. The CpGs best explained by G, G+E or GxE are functionally distinct. The enrichment of GxE variants in GWAS for complex disorders supports their importance for disease risk.

genetics

Genetic networks underlying natural variation in basal and induced activity levels in Drosophila melanogaster

Exercise is recommended by health professionals across the globe as part of a healthy lifestyle to prevent and/or treat the consequences of obesity. While overall, the health benefits of exercise and an active lifestyle are well understood, very little is known about how genetics impacts an individuals inclination for and response to exercise. To address this knowledge gap, we investigated the genetic architecture underlying natural variation in activity levels in the model system Drosophila melanogaster. Activity levels were assayed in the Drosophila Genetics Reference Panel 2 fly strains at baseline and in response to a gentle exercise treatment using the Rotational Exercise Quantification System. We found significant, sex-dependent variation in both activity measures and identified over 100 genes that contribute to basal and induced exercise activity levels. This gene set was enriched for genes with functions in the central nervous system and in neuromuscular junctions and included several candidate genes with known activity phenotypes such as flightlessness or uncoordinated movement. Interestingly, there were also several chromatin proteins among the candidate genes, two of which were validated and shown to impact activity levels. Thus, the study described here reveals the complex genetic architecture controlling basal and exercise-induced activity levels in D. melanogaster and provides a resource for exercise biologists.

genetics

The distribution of deleterious genetic variation in human populations

Population genetic studies suggest that most amino-acid changing mutations are deleterious. Such mutations are of tremendous interest in human population genetics as they are important for the evolutionary process and may contribute risk to common disease. Genomic studies over the past 5 years have documented differences across populations in the number of heterozygous deleterious genotypes, numbers of homozygous derived deleterious genotypes, number of deleterious segregating sites and proportion of sites that are potentially deleterious. These differences have been attributed to population history affecting the ability of natural selection to remove deleterious variants from the population. However, recent studies have suggested that the genetic load may not differ across populations, and that the efficacy of natural selection has not differed across human populations. Here I show that these observations are not incompatible with each other and that the apparent differences are due to examining different features of the genetic data and differing definitions of terms.

Evolutionary Biology

Which genetic variants in DNase I sensitive regions are functional?

Ongoing large experimental characterization is crucial to determine all regulatory sequences, yet we do not know which genetic variants in those regions are non-silent. Here, we present a novel analysis integrating sequence and DNase I footprinting data for 653 samples to predict the impact of a sequence change on transcription factor binding for a panel of 1,372 motifs. Most genetic variants in footprints (5,810,227) do not show evidence of allele-specific binding (ASB). In contrast, functional genetic variants predicted by our computational models are highly enriched for ASB (3,217 SNPs at 20% FDR). Comparing silent to functional non-coding genetic variants, the latter are 1.22-fold enriched for GWAS traits, have lower allele frequencies, and affect footprints more distal to promoters or active in fewer tissues. Finally, integration of the annotations into 18 GWAS meta-studies improves identification of likely causal SNPs and transcription factors relevant for complex traits.

Genomics

C. elegans harbors pervasive cryptic genetic variation for embryogenesis

Conditionally functional mutations are an important class of natural genetic variation, yet little is known about their prevalence in natural populations or their contribution to disease risk. Here, we describe a vast reserve of cryptic genetic variation, alleles that are normally silent but which affect phenotype when the function of other genes is perturbed, in the gene networks of C. elegans embryogenesis. We find evidence that cryptic-effect loci are ubiquitous and segregate at intermediate frequencies in the wild. The cryptic alleles demonstrate low developmental pleiotropy, in that specific, rather than general, perturbations are required to reveal them. Our findings underscore the importance of genetic background in characterizing gene function and provide a model for the expression of conditionally functional effects that may be fundamental in basic mechanisms of trait evolution and the genetic basis of disease susceptibility.

Evolutionary Biology

Demographic inference using genetic data from a single individual: separating population size variation from population structure

The rapid development of sequencing technologies represents new opportunities for population genetics research. It is expected that genomic data will increase our ability to reconstruct the history of populations. While this increase in genetic information will likely help biologists and anthropologists to reconstruct the demographic history of populations, it also represents new challenges. Recent work has shown that structured populations generate signals of population size change. As a consequence it is often difficult to determine whether demographic events such as expansions or contractions (bottlenecks) inferred from genetic data are real or due to the fact that populations are structured in nature. Given that few inferential methods allow us to account for that structure, and that genomic data will necessarily increase the precision of parameter estimates, it is important to develop new approaches. In the present study we analyse two demographic models. The first is a model of instantaneous population size change whereas the second is the classical symmetric island model. We (i) re-derive the distribution of coalescence times under the two models for a sample of size two, (ii) use a maximum likelihood approach to estimate the parameters of these models (iii) validate this estimation procedure under a wide array of parameter combinations, (iv) implement and validate a model choice procedure by using a Kolmogorov-Smirnov test. Altogether we show that it is possible to estimate parameters under several models and perform efficient model choice using genetic data from a single diploid individual.

Evolutionary Biology

A pooling-based approach to mapping genetic variants associated with DNA methylation

DNA methylation is an epigenetic modification that plays a key role in gene regulation. Previous studies have investigated its genetic basis by mapping genetic variants that are associated with DNA methylation at specific sites, but these have been limited to microarrays that cover less than 2% of the genome and cannot account for allele-specific methylation (ASM). Other studies have performed whole-genome bisulfite sequencing on a few individuals, but these lack statistical power to identify variants associated with DNA methylation. We present a novel approach in which bisulfite-treated DNA from many individuals is sequenced together in a single pool, resulting in a truly genome-wide map of DNA methylation. Compared to methods that do not account for ASM, our approach increases statistical power to detect associations while sharply reducing cost, effort, and experimental variability. As a proof of concept, we generated deep sequencing data from a pool of 60 human cell lines; we evaluated almost twice as many CpGs as the largest microarray studies and identified over 2,000 genetic variants associated with DNA methylation. We found that these variants are highly enriched for associations with chromatin accessibility and CTCF binding but are less likely to be associated with traits indirectly linked to DNA, such as gene expression and disease phenotypes. In summary, our approach allows genome-wide mapping of genetic variants associated with DNA methylation in any tissue of any species, without the need for individual-level genotype or methylation data.

Genomics

Using the Phenoscape Knowledgebase to relate genetic perturbations to phenotypic evolution

The abundance of phenotypic diversity among species can enrich our knowledge of development and genetics beyond the limits of variation that can be observed in model organisms. The Phenoscape Knowledgebase (KB) is designed to enable exploration and discovery of phenotypic variation among species. Because phenotypes in the KB are annotated using standard ontologies, evolutionary phenotypes can be compared with phenotypes from genetic perturbations in model organisms. To illustrate the power of this approach, we review the use of the KB to find taxa showing evolutionary variation similar to that of a query gene. Matches are made between the full set of phenotypes described for a gene and an evolutionary profile, the latter of which is defined as the set of phenotypes that are variable among the daughters of any node on the taxonomic tree. Phenoscapes semantic similarity interface allows the user to assess the statistical significance of each match and flags matches that may only result from differences in annotation coverage between genetic and evolutionary studies. Tools such as this will help meet the challenge of relating the growing volume of genetic knowledge in model organisms to the diversity of phenotypes in nature. The Phenoscape KB is available at http://kb.phenoscape.org.

Bioinformatics