Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

The molecular basis of genetic interaction diversity in a metabolic pathway

Our ability to predict the impact of mutations on traits relevant for disease and evolution remains severely limited by the dependence of their effects on the genetic background and environment. Even when molecular interactions between genes are known, it is unclear how these translate to organism-level interactions between alleles. We therefore characterized the interplay of genetic and environmental dependencies in determining fitness by quantifying ~4,000 fitness interactions between expression variants of two metabolic genes, in different environments. We detect a remarkable variety of environment-dependent interactions, and demonstrate they can be quantitatively explained by a mechanistic model accounting for catabolic flux, metabolite toxicity and expression costs. Complex fitness interactions between mutations can therefore be predicted simply from their simultaneous impact on a few connected molecular phenotypes.

genetics

Evidence for secondary-variant genetic burden and non-random distribution across biological modules in a recessive ciliopathy

The influence of genetic background on driver mutations is well established; however, the mechanisms by which the background interacts with Mendelian loci remains unclear. We performed a systematic secondary-variant burden analysis of two independent Bardet-Biedl syndrome (BBS) cohorts with known recessive biallelic pathogenic mutations in one of 17 BBS genes for each individual. We observed a significant enrichment of trans-acting rare nonsynonymous secondary variants compared to either population controls or to a cohort of individuals with a non-BBS diagnosis and recessive variants in the same gene set. Strikingly, we found a significant over-representation of secondary alleles in chaperonin-encoding genes, a finding corroborated by the observation of epistatic interactions involving this complex in vivo. These data indicate a complex genetic architecture for BBS that informs the biological properties of disease modules and presents a model paradigm for secondary-variant burden analysis in recessive disorders.

genetics

Defining the genetic control of human blood plasma N-glycome using genome-wide association study

Glycosylation is a common post-translational modification of proteins. It is known, that glycans are directly involved in the pathophysiology of every major disease. Defining genetic factors altering glycosylation may provide a basis for novel approaches to diagnostic and pharmaceutical applications. Here, we report a genome-wide association study of the human blood plasma N-glycome composition in up to 3811 people. We discovered and replicated twelve loci. This allowed us to demonstrate a clear overlap in genetic control between total plasma and IgG glycosylation. Majority of loci contained genes that encode enzymes directly involved in glycosylation (FUT3/FUT6, FUT8, B3GAT1, ST6GAL1, B4GALT1, ST3GAL4, MGAT3, and MGAT5). We, however, also found loci that are likely to reflect other, more complex, aspects of plasma glycosylation process. Functional genomic annotation suggested the role of DERL3, which potentially highlights the role of glycoprotein degradation pathway, and such transcription factor as IKZF1.

genetics

The Genetic Landscape of Diamond-Blackfan Anemia

Diamond-Blackfan anemia (DBA) is a rare bone marrow failure disorder that affects 1 in 100,000 to 200,000 live births and has been associated with mutations in components of the ribosome. In order to characterize the genetic landscape of this genetically heterogeneous disorder, we recruited a cohort of 472 individuals with a clinical diagnosis of DBA and performed whole exome sequencing (WES). Overall, we identified rare and predicted damaging mutations in likely causal genes for 78% of individuals. The majority of mutations were singletons, absent from population databases, predicted to cause loss of function, and in one of 19 previously reported genes encoding for a diverse set of ribosomal proteins (RPs). Using WES exon coverage estimates, we were able to identify and validate 31 deletions in DBA associated genes. We also observed an enrichment for extended splice site mutations and validated the diverse effects of these mutations using RNA sequencing in patientderived cell lines. Leveraging the size of our cohort, we observed several robust genotype-phenotype associations with congenital abnormalities and treatment outcomes. In addition to comprehensively identifying mutations in known genes, we further identified rare mutations in 7 previously unreported RP genes that may cause DBA. We also identified several distinct disorders that appear to phenocopy DBA, including 9 individuals with biallelic CECR1 mutations that result in deficiency of ADA2. However, no new genes were identified at exome-wide significance, suggesting that there are no unidentified genes containing mutations readily identified by WES that explain > 5% of DBA cases. Overall, this comprehensive report should not only inform clinical practice for DBA patients, but also the design and analysis of future rare variant studies for heterogeneous Mendelian disorders.

genetics

A genetic perspective on Longobard-Era migrations

From the first century AD, Europe has been interested by population movements, commonly known as Barbarian migrations. Among these processes, the one involving the Longobard culture interested a vast region, but its dynamics and demographic impact remains largely unknown. Here we report 87 new complete mitochondrial sequences coming from nine early-medieval cemeteries located along the area interested by the Longobard migration (Czech Republic, Hungary and Italy). From the same locations, we sampled necropolises characterized by cultural markers associated with the Longobard culture (LC) and coeval burials where no such markers were found (NLC). Population genetics analysis and ABC modeling highlighted a similarity between LC individuals, as reflected by a certain degree of genetic continuity between these groups, that reached 70% among Hungary and Italy. Models postulating a contact between LC and NLC communities received also high support, indicating a complex dynamics of admixture in medieval Europe.

genetics

Quantitative approaches to variant classification increase the yield and precision of genetic testing in Mendelian diseases: The case of hypertrophic cardiomyopathy

BackgroundInternational guidelines for variant interpretation in Mendelian disease set stringent criteria to report a variant as (likely) pathogenic, prioritising control of false positive rate over test sensitivity and diagnostic yield. Genetic testing is also more likely informative in individuals with well-characterised variants from extensively studied European-ancestry populations. Inherited cardiomyopathies are relatively common Mendelian diseases that allow empirical calibration and assessment of this framework.\n\nResultsWe compared rare variants in large hypertrophic cardiomyopathy (HCM) cohorts to reference populations to identify variant classes with high prior likelihoods of pathogenicity, as defined by etiological fraction (EF). Analysis of variant distribution identified regions in which variants are significantly enriched in cases and variant location was a better discriminator of pathogenicity than generic computational functional prediction algorithms. Non-truncating variant classes with an EF[≥]0.95, and therefore clinically actionable, were identified in 5 established HCM genes. Applying this approach leads to an estimated 14-20% increase in cases with actionable HCM variants.\n\nConclusionsWhen found in a patient confirmed to have disease, novel variants in some genes and regions are empirically shown to have a sufficiently high probability of pathogenicity to support a \"likely pathogenic\" classification, even without additional segregation or functional data. This could increase the yield of high confidence actionable variants, consistent with the framework and recommendations of current guidelines. The techniques outlined offer a consistent, unbiased and equitable approach to variant interpretation for Mendelian disease genetic testing. We propose adaptations to ACMG/AMP guidelines to incorporate such evidence in a quantitative and transparent manner.

genetics

Measuring intolerance to mutation in human genetics

In numerous applications, from working with animal models to mapping the genetic basis of human disease susceptibility, it is useful to know whether a single disrupting mutation in a gene is likely to be deleterious1-4. With this goal in mind, a number of measures have been developed to identify genes in which protein-truncating variants (PTVs), or other types of mutations, are absent or kept at very low frequency in large population samples--genes that appear \"intolerant to mutation\"3,5-9. One measure in particular, pLI, has been widely adopted7. By contrasting the observed versus expected number of PTVs, it aims to classify genes into three categories, labelled \"null\", \"recessive\" and \"haploinsufficient\"7. Such population genetic approaches can be useful in many applications. As we clarify, however, these measures reflect the strength of selection acting on heterozygotes, and not dominance for fitness or haploinsufficiency for other phenotypes.

genetics

Dissection of complex, fitness-related traits in multiple Drosophila mapping populations offers insight into the genetic control of stress resistance

We leverage two complementary Drosophila melanogaster mapping panels to genetically dissect starvation resistance, an important fitness trait. Using >1600 genotypes of the multiparental Drosophila Synthetic Population Resource (DSPR) we map numerous starvation stress QTL that collectively explain a substantial fraction of trait heritability. QTL effects further allowed us to estimate DSPR founder phenotypes, predictions that were correlated with the actual founder phenotypes. Starvation resistance has been linked to triglyceride level, and while we observe a modest phenotypic correlation between the traits in the DSPR, overlap among the QTL identified for each trait is low. Since we show that DSPR strains with extreme starvation phenotypes also differ in desiccation resistance and activity level, our data imply that multiple physiological mechanisms contribute to starvation variability. We also exploit the Drosophila Genetic Reference Panel (DGRP), identifying a number of sequence variants associated with starvation resistance. Consistent with prior work these sites rarely fall within QTL intervals mapped in the DSPR. Two other groups previously measured starvation resistance in the DGRP, offering a unique opportunity to directly compare mapping results across labs. We found strong phenotypic correlations among studies, but extremely low overlap in the sets of genomewide significant sites. Despite this, our analyses revealed that the most highly-associated variants from each study typically showed the same additive effect sign in independent studies, in contrast to otherwise equivalent sets of random variants. This consistency provides evidence for reproducible trait-associated sites in a widely-used mapping panel, and highlights the polygenic nature of starvation resistance.

genetics

Efficient implementation of penalized regression for genetic risk prediction

Polygenic Risk Scores (PRS) consist in combining the information across many single-nucleotide polymorphisms (SNPs) in a score reflecting the genetic risk of developing a disease. PRS might have a major impact on public health, possibly allowing for screening campaigns to identify high-genetic risk individuals for a given disease. The \"Clumping+Thresholding\" (C+T) approach is the most common method to derive PRS. C+T uses only univariate genome-wide association studies (GWAS) summary statistics, which makes it fast and easy to use. However, previous work showed that jointly estimating SNP effects for computing PRS has the potential to significantly improve the predictive performance of PRS as compared to C+T.\n\nIn this paper, we present an efficient method to jointly estimate SNP effects, allowing for practical application of penalized logistic regression (PLR) on modern datasets including hundreds of thousands of individuals. Moreover, our implementation of PLR directly includes automatic choices for hyper-parameters. The choice of hyper-parameters for a predictive model is very important since it can dramatically impact its predictive performance. As an example, AUC values range from less than 60% to 90% in a model with 30 causal SNPs, depending on the p-value threshold in C+T.\n\nWe compare the performance of PLR, C+T and a derivation of random forests using both real and simulated data. PLR consistently achieves higher predictive performance than the two other methods while being as fast as C+T. We find that improvement in predictive performance is more pronounced when there are few effects located in nearby genomic regions with correlated SNPs; for instance, AUC values increase from 83% with the best prediction of C+T to 92.5% with PLR. We confirm these results in a data analysis of a case-control study for celiac disease where PLR and the standard C+T method achieve AUC of 89% and of 82.5%.\n\nIn conclusion, our study demonstrates that penalized logistic regression can achieve more discriminative polygenic risk scores, while being applicable to large-scale individual-level data thanks to the implementation we provide in the R package bigstatsr.

genetics

Estimation of realized rates of genetic gain and indicators for breeding program assessment

Routine estimation of the rate of genetic gain ({Delta}Gt) realized by a breeding program has been proposed as a means to monitor its effectiveness. Several methods of realized{Delta} Gt estimation have been utilized in other studies, but none have been objectively evaluated in a plant breeding context. Stochastic simulations of 80 rice (Oryza sativa) breeding programs over 28 years were done to generate data used to evaluate five methods of realized{Delta} Gt estimation in terms of error, precision, efficiency and correlation between true and predicted annual mean breeding values. Two indicators of{Delta} Gt, the expected{Delta} Gt and the average number of equivalent complete generations (EqCg), were described and evaluated. At best, estimates of realized{Delta} Gt were over or underestimated by 15% and 27% when considering all 28 years and the past 15 years of breeding respectively. The best methods were the control population, estimated breeding value, and ERA trial methods. Among these, correlations between true and estimated{Delta} Gt were at best 0.59, indicating that these methods cannot very accurately rank breeding programs in terms of realized{Delta} Gt. The expected{Delta} Gt and the average EqCg were shown to be useful indicators for determining if a non-zero genetic gain is expected. Determining which of the three best realized{Delta} Gt estimation methods evaluated, if any, would be appropriate for any given breeding program should be done with careful consideration of the objectives, resources, seed stocks, and structure of the data available.

genetics

Exploring Genetic Variation That Influences Brain Methylation In Attention-Deficit/Hyperactivity Disorder

Attention-deficit/hyperactivity disorder (ADHD) is a neurodevelopmental disorder caused by an interplay of genetic and environmental factors. Epigenetics is crucial to lasting changes in gene expression in the brain. Recent studies suggest a role for DNA methylation in ADHD. We explored the contribution to ADHD of allele-specific methylation (ASM), an epigenetic mechanism that involves SNPs correlating with differential levels of DNA methylation at CpG sites. We selected 3,896 tagSNPs reported to influence methylation in human brain regions and performed a case-control association study using the summary statistics from the largest GWAS meta-analysis of ADHD, comprising 20,183 cases and 35,191 controls. We identified associations with eight tagSNPs that were significant at a 5% False Discovery Rate (FDR). These SNPs correlated with methylation of CpG sites lying in the promoter regions of six genes. Since methylation may affect gene expression, we inspected these ASM SNPs together with 52 ASM SNPs in high LD with them for eQTLs in brain tissues and observed that the expression of three of those genes was affected by them. ADHD risk alleles correlated with increased expression (and decreased methylation) of ARTN and PIDD1 and with a decreased expression (and increased methylation) of C2orf82. Furthermore, these three genes were predicted to have altered expression in ADHD, and genetic variants in C2orf82 correlated with brain volumes. In summary, we followed a systematic approach to identify risk variants for ADHD that correlated with differential cis-methylation, identifying three novel genes contributing to the disorder.

genetics

Quantifying Heterogeneity in the Genetic Architecture of Complex Traits Between Ethnically Diverse Groups using Random Effect Interaction Models

In humans, most genome-wide association studies have been conducted using data from Caucasians and many of the reported findings have not replicated in other populations. This lack of replication may be due to statistical issues (small sample size, confounding) or perhaps more fundamentally to differences in the genetic architecture of traits between ethnically diverse subpopulations. What aspects of the genetic architecture of traits vary between subpopulations and how can this be quantified? We consider studying effect heterogeneity using random-effect Bayesian interaction models. The proposed methodology can be applied using shrinkage and variable selection methods and produces useful information about effect heterogeneity in the form of whole-genome summaries (e.g., SNP-heritability and the average correlation of effects) as well as SNP-specific attributes. Using simulations, we show that the proposed methodology yields (nearly) unbiased estimates of genomic heritability and of the average correlation of effects between groups when the sample size is not too small relative to the number of SNPs used. Subsequently, we used the proposed methodology for the analyses of four complex human traits (standing height, high-density lipoprotein, low-density lipoprotein, and serum urate levels) in European-Americans (EAs) and African-Americans (AAs). The estimated correlations of effects between the two subpopulations was well below unity for all the traits, ranging from 0.73 to 0.50. The extent of effect heterogeneity varied between traits and SNP-sets. Height showed less differences in SNP effects between AAs and EAs whereas HDL, a trait highly influenced by life-style, exhibited greater extent of effect heterogeneity. For all the traits we observed substantial variability in effect heterogeneity across SNPs, suggesting it varies between regions of the genome.

genetics

The genetic legacy of continental scale admixture in Indian Austroasiatic speakers

Surrounded by speakers of Indo-European, Dravidian and Tibeto-Burman languages, around 11 million Munda (a branch of Austroasiatic language family) speakers live in the densely populated and genetically diverse South Asia. Their genetic makeup holds components characteristic of South Asians as well as Southeast Asians. The admixture time between these components has been previously estimated on the basis of archaeology, linguistics and uniparental markers. Using genome-wide genotype data of 102 Munda speakers and contextual data from South and Southeast Asia, we retrieved admixture dates between 2000 - 3800 years ago for different populations of Munda. The best modern proxies for the source populations for the admixture with proportions 0.78/0.22 are Lao people from Laos and Dravidian speakers from Kerala in India, while the South Asian population(s), with whom the incoming Southeast Asians intermixed, had a smaller proportion of West Eurasian component than contemporary proxies. Somewhat surprisingly Malaysian Peninsular tribes rather than the geographically closer Austroasiatic languages speakers like Vietnamese and Cambodians show highest sharing of IBD segments with the Munda. In addition, we affirmed that the grouping of the Munda speakers into North and South Munda based on linguistics is in concordance with genome-wide data.

genetics

Survival protection mechanisms and genetic variability induction after stress: two sides of the same Hsp70 coin

Previous studies have shown that heat shock stress may increase transcription levels and, in some cases, also the transposition of certain transposable elements (TEs) in Drosophila and other organisms. Other studies have also demonstrated that heat shock chaperones as Hsp90 and Hop are involved in repressing transposons activity in Drosophila melanogaster by their involvement in crucial steps of the biogenesis of Piwi-interacting RNAs (piRNAs), the largest class of germline-enriched small non-coding RNA implicated in the epigenetic silencing of TEs. However, a satisfying picture of how many chaperones and their respective functional roles could be involved in repressing transposons in germ cell is still unknown. Here we show that in Drosophila heat shock activates transposon's expression at post-transcriptional level by disrupting a repressive chaperone complex by a decisive role of the stress-inducible chaperone Hsp70. We found that stress-induced transposons activation is triggered by an interaction of Hsp70 with the Hsc70-Hsp90 complex and other factors all involved in piRNA biogenesis in both ovaries and testes. Such interaction induces a displacement of all such factors to the lysosomes resulting in a functional collapse of piRNA biogenesis. In support of a significant role of Hsp70 in transposon activation after stress, we found that the expression under normal conditions of Hsp70 in transgenic flies increases the amount of transposon transcripts and displaces the components of chaperon machinery outside the nuage as observed after heat shock. So that, our results demonstrate that heat shock stress is capable to increase the expression of transposons at post-transcriptional level by affecting piRNA biogenesis through the action of the inducible chaperone Hsp70. We think that such mechanism proposes relevant evolutionary implications. In presence of drastic environmental changes, Hsp70 plays a key dual role in increasing both the survival probability of individuals and the genetic variability in their germ cells. This in turn should be translated into an increase of genetic variability inside the populations thus potentiating their evolutionary plasticity and evolvability.

genetics

Identifying novel subtypes of irritability using a developmental genetic approach

ObjectiveIrritability is a common reason for referral to services, strongly associated with impairment and negative outcomes, but is a nosological and treatment challenge. A major issue is how irritability should be conceptualized. This study used a developmental approach to test the hypothesis that there are several forms of irritability, including a neurodevelopmental/ADHD-like subtype with onset in childhood and a depression/mood subtype with onset in adolescence.\n\nMethodData were analyzed in the Avon Longitudinal Study of Parents and Children, a prospective UK population-based cohort. Irritability trajectory-classes were estimated for 7924 individuals with data at multiple time-points across childhood and adolescence (4 possible time-points from approximately ages 7 to 15 years). Psychiatric diagnoses were assessed at approximately ages 7 and 15 years. Psychiatric genetic risk was indexed by polygenic risk scores (PRS) for attention-deficit/hyperactivity disorder (ADHD) and major depressive disorder (MDD) derived using large genome-wide association study results.\n\nResultsFive irritability trajectory classes were identified: low (81.2%), decreasing (5.6%), increasing (5.5%), late-childhood limited (5.2%) and high-persistent (2.4%). The early-onset, high-persistent trajectory was associated with male preponderance, childhood ADHD (OR=108.64 (57.45-204.41), p<0.001) and ADHD PRS (OR=1.31 (1.09-1.58), p=0.005); the adolescent-onset, increasing trajectory was associated with female preponderance, adolescent MDD (OR=5.14 (2.47-10.73), p<0.001) and MDD PRS (OR=1.20, (1.05-1.38), p=0.009). Both trajectory classes were associated with MDD diagnosis and ADHD genetic risk.\n\nConclusionsThe developmental context of irritability may be important in its conceptualization: early-onset persistent irritability maybe more neurodevelopmental/ADHD-like and later-onset irritability more depression/mood-like. This has implications for treatment as well as nosology.

genetics

Applicability of the mutation-selection balance model to population genetics of heterozygous protein-truncating variants in humans

The fate of alleles in the human population is believed to be highly affected by the stochastic force of genetic drift. Estimation of the strength of natural selection in humans generally necessitates a careful modeling of drift including complex effects of the population history and structure. Protein truncating variants (PTVs) are expected to evolve under strong purifying selection and to have a relatively high per-gene mutation rate. Thus, it is appealing to model the population genetics of PTVs under a simple deterministic mutation-selection balance, as has been proposed earlier [1]. Here, we investigated the limits of this approximation using both computer simulations and data-driven approaches. Our simulations rely on a model of demographic history estimated from 33,370 individual exomes of the Non-Finnish European subset of the ExAC dataset [2]. Additionally, we compared the African and European subset of the ExAC study and analyzed de novo PTVs. We show that the mutation-selection balance model is applicable to the majority of human genes, but not to genes under the weakest selection.

genetics

The molecular genetics of hand preference revisited

Hand preference is a prominent behavioural trait linked to human brain asymmetry. A handful of genetic variants have been reported to associate with hand preference or quantitative measures related to it. Most of these reports were on the basis of limited sample sizes, by current standards for genetic analysis of complex traits. Here we performed a genome-wide association analysis of hand preference in the large, population-based UK Biobank cohort (N=331,037). We used gene-set enrichment analysis to investigate whether genes involved in visceral asymmetry are particularly relevant to hand preference, following one previous report. We found no evidence implicating any specific candidate variants previously reported. We also found no evidence that genes involved in visceral laterality play a role in hand preference. It remains possible that some of the previously reported genes or pathways are relevant to hand preference as assessed in other ways, or else are relevant within specific disorder populations. However, some or all of the earlier findings are likely to be false positives, and none of them appear relevant to hand preference as defined categorically in the general population. Within the UK Biobank itself, a significant association implicates the gene MAP2 in handedness.

genetics

Genetics of single-cell protein abundance variation in large yeast populations

Many DNA sequence variants influence phenotypes by altering gene expression. Our understanding of these variants is limited by sample sizes of current studies and by measurements of mRNA rather than protein abundance. We developed a powerful method for identifying genetic loci that influence protein expression in very large populations of the yeast Saccharomyes cerevisiae. The method measures single-cell protein abundance through the use of green-fluorescent-protein tags. We applied this method to 160 genes and detected many more loci per gene than previous studies. We also observed closer correspondence between loci that influence protein abundance and loci that influence mRNA abundance of a given gene. Most loci cluster at hotspot locations that influence multiple proteins--in some cases, more than half of those examined. The variants that underlie these hotspots have profound effects on the gene regulatory network and provide insights into genetic variation in cell physiology between yeast strains.

Genomics