Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Transcriptome analysis of genetically matched human induced pluripotent stem cells disomic or trisomic for chromosome 21.

Trisomy of chromosome 21, the genetic cause of Down syndrome, has the potential to alter expression of genes on chromosome 21, as well as other locations throughout the genome. These transcriptome changes are likely to underlie the Down syndrome clinical phenotypes. We have employed RNA-seq to undertake an in-depth analysis of transcriptome changes resulting from trisomy of chromosome 21, using induced pluripotent stem cells (iPSCs) derived from a single individual with Down syndrome. These cells were originally derived by Li et al, who genetically targeted chromosome 21 in trisomic iPSCs, allowing selection of disomic sibling iPSC clones. Analyses were conducted on trisomic/disomic cell pairs maintained as iPSCs or differentiated into cortical neuronal cultures. In addition to characterization of gene expression levels, we have also investigated patterns of RNA adenosine-to-inosine editing, alternative splicing, and repetitive element expression, aspects of the transcriptome that have not been significantly characterized in the context of Down syndrome. We identified significant changes in transcript accumulation associated with chromosome 21 trisomy, as well as changes in alternative splicing and repetitive element transcripts. Unexpectedly, the trisomic iPSCs we characterized expressed higher levels of neuronal transcripts than control disomic iPSCs, and readily differentiated into cortical neurons, in contrast to another reported study. Comparison of our transcriptome data with similar studies of trisomic iPSCs suggests that trisomy of chromosome 21 may not intrinsically limit neuronal differentiation, but instead may interfere with the maintenance of pluripotency.

genetics

Genetic basis of melanin pigmentation in butterfly wings

Despite the variety, prominence, and adaptive significance of butterfly wing patterns surprisingly little known about the genetic basis of wing color diversity. Even though there is intense interest in wing pattern evolution and development, the technical challenge of genetically manipulating butterflies has slowed efforts to functionally characterize color pattern development genes. To identify candidate wing pigmentation genes we used RNA-seq to characterize transcription across multiple stages of butterfly wing development, and between different color pattern elements, in the painted lady butterfly Vanessa cardui. This allowed us to pinpoint genes specifically associated with red and black pigment patterns. To test the functions of a subset of genes associated with presumptive melanin pigmentation we used CRISPR/Cas9 genome editing in four different butterfly genera. pale, Ddc, and yellow knockouts displayed reduction of melanin pigmentation, consistent with previous findings in other insects. Interestingly, however, yellow-d, ebony, and black knockouts revealed that these genes have localized effects on tuning the color of red, brown, and ochre pattern elements. These results point to previously undescribed mechanisms for modulating the color of specific wing pattern elements in butterflies, and provide an expanded portrait of the insect melanin pathway.

genetics

The Genetic History of Northern Europe

Recent ancient DNA studies have revealed that the genetic history of modern Europeans was shaped by a series of migration and admixture events between deeply diverged groups. While these events are well described in Central and Southern Europe, genetic evidence from Northern Europe surrounding the Baltic Sea is still sparse. Here we report genome-wide DNA data from 24 ancient North Europeans ranging from [~]7,500 to 200 calBCE spanning the transition from a hunter-gatherer to an agricultural lifestyle, as well as the adoption of bronze metallurgy. We show that Scandinavia was settled after the retreat of the glacial ice sheets from a southern and a northern route, and that the first Scandinavian Neolithic farmers derive their ancestry from Anatolia 1000 years earlier than previously demonstrated. The range of Western European Mesolithic hunter-gatherers extended to the east of the Baltic Sea, where these populations persisted without gene-flow from Central European farmers until around 2,900 calBCE when the arrival of steppe pastoralists introduced a major shift in economy and established wide-reaching networks of contact within the Corded Ware Complex.

genetics

Genotype-phenotype association mining in bipolar disorder: market research meets complex genetics

Disentangling the etiology of common, complex diseases is a major challenge in genetic research. For bipolar disorder (BD), several genome-wide association studies (GWAS) have been performed. Similar to other complex disorders, major breakthroughs in explaining the high heritability of BD through GWAS have remained elusive. To overcome this dilemma, genetic research into BD, has embraced a variety of strategies such as the formation of large consortia to increase sample size and sequencing approaches. Here we advocate a complementary approach making use of already existing GWAS data: applying a data mining procedure to identify yet undetected genotype-phenotype relationships. We adapted association rule mining, a data mining technique traditionally used in retail market research, to identify frequent and characteristic genotype patterns showing strong associations to phenotype clusters. We applied this strategy to three independent GWAS datasets from 2,835 phenotypically characterized patients with BD. In a discovery step, 20,882 candidate association rules were extracted. Two of these - one associated with eating disorder and the other with anxiety - remained significant in an independent dataset after robust correction for multiple testing, showing considerable effect sizes (odds ratio ~ 3.4 and 3.0, respectively). Our approach may help detect novel specific genotype-phenotype relationships in BD typically not explored by analyses like GWAS. While we adapted the data mining tool within the context of BD gene discovery, it may facilitate identifying highly specific genotype-phenotype relationships in subsets of genome-wide data sets of other complex phenotype with similar epidemiological properties and challenges to gene discovery efforts.

genomics

Genetic control of age-related gene expression and complex traits in the human brain

Age is the primary risk factor for many of the most common human diseases--particularly neurodegenerative diseases--yet we currently have a very limited understanding of how each individuals genome affects the aging process. Here we introduce a method to map genetic variants associated with age-related gene expression patterns, which we call temporal expression quantitative trait loci (teQTL). We found that these loci are markedly enriched in the human brain and are associated with neurodegenerative diseases such as Alzheimers disease and Creutzfeldt-Jakob disease. Examining potential molecular mechanisms, we found that age-related changes in DNA methylation can explain some cis-acting teQTLs, and that trans-acting teQTLs can be mediated by microRNAs. Our results suggest that genetic variants modifying age-related patterns of gene expression, acting through both cis- and trans-acting molecular mechanisms, could play a role in the pathogenesis of diverse neurological diseases.

genetics

Natural Genetic Variation Can Independently Tune The Induced Fraction And Induction Level Of A Bimodal Signaling Response

Bimodal gene expression by genetically identical cells is a pervasive feature of signaling networks. In the galactose-utilization (GAL) pathway of Saccharomyces cerevisiae, induction can be unimodal or bimodal depending on natural genetic variation and pre-induction conditions. Here, we find that this variation of modality is regulated by an interplay between two features of the pathway response, the fraction of cells that are in the induced subpopulation and their expression level. Combined, the variations in these features are sufficient to explain the observed effects of natural variation and pre-induction conditions on the modality of induction in both mechanistic and phenomenological models. Both natural variation and pre-induction conditions act by modulating the expression and function of the galactose sensor GAL3. The ability to alter modality may allow organisms to adapt their level of "bet hedging" to the conditions they experience, and thus help optimize fitness in complex, fluctuating natural environments.

genetics

Genetic Instrumental Variable (GIV) Regression: Explaining Socioeconomic And Health Outcomes In Non-Experimental Data

Identifying causal effects in non-experimental data is an enduring challenge. One proposed solution that recently gained popularity is the idea to use genes as instrumental variables (i.e. Mendelian Randomization - MR). However, this approach is problematic because many variables of interest are genetically correlated, which implies the possibility that many genes could affect both the exposure and the outcome directly or via unobserved confounding factors. Thus, pleiotropic effects of genes are themselves a source of bias in non-experimental data that would also undermine the ability of MR to correct for endogeneity bias from non-genetic sources. Here, we propose an alternative approach, GIV regression, that provides estimates for the effect of an exposure on an outcome in the presence of pleiotropy. As a valuable byproduct, GIV regression also provides accurate estimates of the chip heritability of the outcome variable. GIV regression uses polygenic scores (PGS) for the outcome of interest which can be constructed from genome-wide association study (GWAS) results. By splitting the GWAS sample for the outcome into non-overlapping subsamples, we obtain multiple indicators of the outcome PGS that can be used as instruments for each other, and, in combination with other methods such as sibling fixed effects, can address endogeneity bias from both pleiotropy and the environment. In two empirical applications, we demonstrate that our approach produces reasonable estimates of the chip heritability of educational attainment (EA) and show that standard regression and MR provide upwardly biased estimates of the effect of body height on EA.

genetics

The Stratification Of Major Depressive Disorder Into Genetic Subgroups

Depression is a common and clinically heterogeneous mental health disorder that is frequently comorbid with other diseases and conditions. Stratification of depression may align sub-diagnoses more closely with their underling aetiology and provide more tractable targets for research and effective treatment. In the current study, we investigated whether genetic data could be used to identify subgroups within people with depression using the UK Biobank. Examination of cross-locus correlations was used to test for evidence of subgroups by examining whether there was clustering of independent genetic variants associated with eleven other complex traits and disorders in people with depression. We found evidence of a subgroup within depression using age of natural menopause variants (P = 1.69 x 10-3) and this effect remained significant in females (P = 1.18 x 10-3), but not males (P = 0.186). However, no evidence for this subgroup (P > 0.05) was found in Generation Scotland, iPSYCH, a UK Biobank replication cohort or the GERA cohort. In the UK Biobank, having depression was also associated with a later age of menopause (beta = 0.34, standard error = 0.06, P = 9.92 x 10-8). A potential age of natural menopause subgroup within depression and the association between depression and a later age of menopause suggests that they partially share a developmental pathway.

genetics

Genome-Wide Association Study Reveals Genetic Link Between Diarrhea-Associated Entamoeba histolytica Infection And Inflammatory Bowel Disease

Diarrhea is the second leading cause of death for children globally, causing 760,000 deaths each year in children under the age of 5. Amoebic dysentery contributes significantly to this burden, especially in developing countries. We hypothesize that genetic variation contributes to susceptibility to diarrhea-associated Entamoeba histolytica infection in Bangladeshi infants; thus, we conducted a genome-wide association study (GWAS) in two independent birth cohorts of diarrhea-associated E. histolytica infection. Cases were defined as children with at least one diarrheal episode positive for E. histolytica through either PCR or ELISA within the first year of life. Controls were children without any episodes positive for E. histolytica in the same time frame. Meta-analyses under a fixed-effects inverse variance weighting model identified variants in two neighboring genes on chromosome 10: CUL2 (cullin 2) and CREM (cAMP responsive element modulator) associated with E. histolytica infection, with SNP rs58000832 achieving genome-wide significance (Pmeta=4.2x10-10). Each additional risk allele (an intergenic insertion between CREM and CCNY) of rs58000832 conferred 2.5 increased odds of a diarrhea-associated E. histolytica infection. The most associated SNP within a gene was in an intron of CREM (rs58468685, Pmeta=2.3x10-9), which with CUL2, has been implicated as a susceptibility locus for Inflammatory Bowel Disease (IBD) and Crohns Disease. Gene expression resources suggest these loci are related to the higher expression of CREM, but not CUL2. Increased CREM expression is also observed in early E. histolytica infection. Further, CREM-/- mice were more susceptible to E. histolytica amebic colitis. These genetic associations reinforce the pathological similarities observed in gut inflammation between E. histolytica infection and IBD.

genetics

The iPSYCH2012 case-cohort sample: New directions for unravelling genetic and environmental architectures of severe mental disorders

The iPSYCH consortium has established a large Danish population-based Case-Cohort sample (iPSYCH2012) aimed at unravelling the genetic and environmental architecture of severe mental disorders. The iPSYCH2012 sample is nested within the entire Danish population born 1981-2005 including 1,472,762 persons. This paper introduces the iPSYCH2012 sample and outlines key future research directions. Cases were identified as persons with schizophrenia (N=3,540), autism (N=16,146), ADHD (N=18,726), and affective disorder (N=26,380), of which 1928 had bipolar affective disorder. Controls were randomly sampled individuals (N=30,000). Within the sample of 86,189 individuals, a total of 57,377 individuals had at least one major mental disorder. DNA was extracted from the neonatal dried blood spot samples obtained from the Danish Neonatal Screening Biobank and genotyped using the Illumina PsychChip. Genotyping was successful for 90% of the sample. The assessments of exome sequencing, methylation profiling, metabolome profiling, vitamin-D, inflammatory and neurotrophic factors are in progress. For each individual, the iPSYCH2012 sample also includes longitudinal information on health, prescribed medicine, social and socioeconomic information and analogous information among relatives. To the best of our knowledge, the iPSYCH2012 sample is the largest and most comprehensive data source for the combined study of genetic and environmental aetiologies of severe mental disorders.

genetics

Comprehensive genome and transcriptome analysis reveals genetic basis for gene fusions in cancer

Gene fusions are an important class of cancer-driving events with therapeutic and diagnostic values, yet their underlying genetic mechanisms have not been systematically characterized. Here by combining RNA and whole genome DNA sequencing data from 1188 donors across 27 cancer types we obtained a list of 3297 high-confidence tumour-specific gene fusions, 82% of which had structural variant (SV) support and 2372 of which were novel. Such a large collection of RNA and DNA alterations provides the first opportunity to systematically classify the gene fusions at a mechanistic level. While many could be explained by single SVs, numerous fusions involved series of structural rearrangements and thus are composite fusions. We discovered 75 fusions of a novel class of inter-chromosomal composite fusions, termed bridged fusions, in which a third genomic location bridged two different genes. In addition, we identified 522 fusions involving non-coding genes and 157 ORF-retaining fusions, in which the complete open reading frame of one gene was fused to the UTR region of another. Although only a small proportion (5%) of the discovered fusions were recurrent, we found a set of highly recurrent fusion partner genes, which exhibited strong 5 or 3 bias and were significantly enriched for cancer genes. Our findings broaden the view of the gene fusion landscape and reveal the general properties of genetic alterations underlying gene fusions for the first time.

genetics

Genome-wide analysis reveals distinct genetic mechanisms of diet-dependent lifespan and healthspan in D. melanogaster

Dietary restriction (DR) robustly extends lifespan and delays age-related diseases across species. An underlying assumption in aging research has been that DR mimetics extend both lifespan and healthspan jointly, though this has not been rigorously tested in different genetic backgrounds. Furthermore, nutrient response genes important for lifespan or healthspan extension remain underexplored, especially in natural populations. To address these gaps, we utilized over 150 DGRP strains to measure nutrient-dependent changes in lifespan and age-related climbing ability to measure healthspan. DR extended lifespan and delayed decline in climbing ability on average, but there was no evidence of correlation between these traits across individual strains. Through GWAS, we then identified and validated jughead and Ferredoxin as determinants of diet-dependent lifespan, and Daedalus for diet-dependent physical activity. Modulating these genes produced independent effects on lifespan and climbing ability, further suggesting that these age-related traits are likely to be regulated through distinct genetic mechanisms.

genetics

A genetic investigation of sex bias in the prevalence of attention deficit hyperactivity disorder

Attention-deficit/hyperactivity disorder (ADHD) shows substantial heritability and is 2-7 times more common in males than females. We examined two putative genetic mechanisms underlying this sex bias: sex-specific heterogeneity and higher burden of risk in female cases. We analyzed genome-wide common variants from the Psychiatric Genomics Consortium and iPSYCH Project (20,183 cases, 35,191 controls) and Swedish population-registry data (N=77,905 cases, N=1,874,637 population controls). We find strong genetic correlation for ADHD across sex and no mean difference in polygenic burden across sex. In contrast, siblings of female probands are at an increased risk of ADHD, compared to siblings of male probands. The results also suggest that females with ADHD are at especially high risk of comorbid developmental conditions. Overall, this study supports a greater familial burden of risk in females with ADHD and some clinical and etiological heterogeneity. However, autosomal common variants largely do not explain the sex bias in ADHD prevalence.

genetics

Psychosis and the level of mood incongruence in Bipolar Disorder are related to genetic liability for Schizophrenia

ImportanceBipolar disorder (BD) overlaps schizophrenia in its clinical presentation and genetic liability. Alternative approaches to patient stratification beyond current diagnostic categories are needed to understand the underlying disease processes/mechanisms.\n\nObjectivesTo investigate the relationship between common-variant liability for schizophrenia, indexed by polygenic risk scores (PRS) and psychotic presentations of BD, using clinical descriptions which consider both occurrence and level of mood-incongruent psychotic features.\n\nDesignCase-control design: using multinomial logistic regression, to estimate differential associations of PRS across categories of cases and controls.\n\nSettings & Participants4399 BDcases, mean [sd] age-at-interview 46[12] years, of which 2966 were woman (67%) from the BD Research Network (BDRN) were included in the final analyses, with data for 4976 schizophrenia cases and 9012 controls from the Type-1 diabetes genetics consortium and Generation Scotland included for comparison.\n\nExposureStandardised PRS, calculated using alleles with an association p-value threshold < 0.05 in the second Psychiatric Genomics Consortium genome-wide association study of schizophrenia, adjusted for the first 10 population principal components and genotyping-platform.\n\nMain outcome measureMultinomial logit models estimated PRS associations with BD stratified by (1) Research Diagnostic Criteria (RDC) BD subtypes (2) Lifetime occurrence of psychosis.(3) Lifetime mood-incongruent psychotic features and (4) ordinal logistic regression examined PRS associations across levels of mood-incongruence. Ratings were derived from the Schedule for Clinical Assessment in Neuropsychiatry interview (SCAN) and the Bipolar Affective Disorder Dimension Scale (BADDS).\n\nResultsAcross clinical phenotypes, there was an exposure-response gradient with the strongest PRS association for schizophrenia (RR=1.94, (95% C.1.1.86, 2.01)), then schizoaffective BD (RR=1.37, (95% C.I. 1.22, 1.54)), BD I (RR= 1.30, (95% C.I. 1.24, 1.36)) and BD II (RR=1.04, (95% C.1. 0.97, 1.11)). Within BD cases, there was an effect gradient, indexed by the nature of psychosis, with prominent mood-incongruent psychotic features having the strongest association (RR=1.46, (95% C.1.1.36, 1.57)), followed by mood-congruent psychosis (RR= 1.24, (95% C.1. 1.17, 1.33)) and lastly, BD cases with no history of psychosis (RR= 1.09, (95% C.1. 1.04, 1.15)).\n\nConclusionWe show for the first time a polygenic-risk gradient, across schizophrenia and bipolar disorder, indexed by the occurrence and level of mood-incongruent psychotic symptoms.

genetics

Forward Genetic ENU Mutagenesis Screen for Mouse Models of Chronic Fatigue Identifies a Novel Mutation in Slc2a4 (GLUT4)

In a screen of voluntary wheel-running behavior designed to identify genetic mouse models of chronic fatigue in ENU mutagenized C57BL/6J mice, we discovered two lines that showed aberrant wheel-running patterns. These lines both stem from a single original founder identified as a low body-weight candidate in a recessive screen. Progeny from both of these lines showed the abnormal wheel-running behavior, with affected mice showing significantly lower daily activity levels than unaffected mice. They also exhibited low amplitude circadian rhythms, consisting of lower activity levels during the normal active phase, and increased levels of activity during the rest or light phase, but only a modest alteration in free-running period. Their activity is not consolidated into longer bouts, but is frequently interrupted with periods of inactivity throughout the dark phase of the light-dark (LD) cycle. As seen with the low body weight, expression of the behavioral phenotypes in offspring of strategic crosses was consistent with a recessive heritance pattern. Mapping of these phenotypic abnormalities showed linkage to a single locus on chromosome 11, and whole exome sequencing (WES) identified a single point mutation in the Slc2a4 gene encoding the GLUT4 insulin-responsive glucose transporter. The single nucleotide change (A to T) was found in the distal end of exon 10, and results in a premature stop (Y440*). To our knowledge, this is the first time a mutation in this gene has been shown to result in extensive changes in general behavioral patterns.\n\nSIGNIFICANCE STATEMENTChronic fatigue is a debilitating and devastating disorder with widespread consequences for both the patient and the persons around them, but effective treatment strategies are lacking. The identification of novel genetic mouse models of chronic fatigue may prove invaluable for the study of its underlying physiological mechanisms and for the testing of treatments and interventions. A novel mutation in Slc2a4 (GLUT4) was identified in a forward mutagenesis screen because affected mice showed abnormal daily patterns and levels of wheel running consistent with chronic fatigue. This new mouse model may shed light on the pathophysiology of chronic fatigue.

genetics

A genetic-based algorithm for recovery: A pilot study

Exercise training creates a number of physical challenges to the body, the overcoming of which drives exercise adaptation. The balance between sufficient stress and recovery is a crucial, but often under-explored, area within exercise training. Genetic variation can also predispose some individuals to a greater need for recovery after exercise. In this pilot study, 18 male soccer players underwent a repeated sprint training session. Countermovement jump (CMJ) heights were recorded immediately pre-and post-training, and at 24-and 48-hours post-training. The reduction in CMJ height was greatest at all post-training time points in subjects with a larger number of gene variants associated with a reduced exercise recovery. This suggests that knowledge of genetic information can be important in individualizing recovery timings and modalities in athletes following training.

genetics

Genetics of vaccination-related narcolepsy

Narcolepsy type 1 is a severe hypersomnia affecting 1/3000 individuals. It is caused by a loss of neurons producing hypocretin/orexin in the hypothalamus. In 2009/2010, an immunization campaign directed towards the new pandemic H1N1 Influenza-A strain was launched and increased risk of narcolepsy reported in Northern European countries following vaccination with Pandemrix(R), an adjuvanted H1N1 vaccine resulting in ~250 vaccination-related cases in Finland alone. Using whole genome sequencing data of 2000 controls, exome sequencing data of 5000 controls and HumanCoreExome chip genotypes of 81 cases with vaccination-related narcolepsy and 2796 controls, we, built a multilocus genetic risk score with established narcolepsy risk variants. We also analyzed, whether novel risk variants would explain vaccine-related narcolepsy. We found that previously discovered risk variants had strong predictive power (accuracy of 73% and P<2.2*10-16; and ROC curve AUC 0.88) in vaccine-related narcolepsy cases with only 4.9% of cases being assigned to the low risk category. Our findings indicate genetic predisposition to vaccine-triggered narcolepsy, with the possibility of identifying 95% of people at risk.

genetics

Long-read sequencing across the C9orf72 ‘GGGGCC’ repeat expansion: implications for genetic discovery efforts in human disease

Background: Many neurodegenerative diseases are caused by nucleotide repeat expansions, but most expansions, like the C9orf72 GGGGCC (G4C2) repeat that causes approximately 5-7% of all amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD) cases, are too long to sequence using short-read sequencing technologies. It is unclear whether long-read sequencing technologies can traverse these long, challenging repeat expansions. Here, we demonstrate that two long-read sequencing technologies, Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), can sequence through disease-causing repeats cloned into plasmids, including the FTD/ALS-causing G4C2 repeat expansion. We also report the first long-read sequencing data characterizing the C9orf72 G4C2 repeat expansion at the nucleotide level in two symptomatic expansion carriers using PacBio whole-genome sequencing and a no-amplification (No-Amp) targeted approach based on CRISPR/Cas9.\n\nResults: Both the PacBio and ONT platforms successfully sequenced through the repeat expansions in plasmids. Throughput on the MinlON was a challenge for whole-genome sequencing; we were unable to attain reads covering the human C9orf72 repeat expansion using 15 flow cells. We obtained 8x coverage across the C9orf72 locus using the PacBio Sequel, accurately reporting the unexpanded allele at eight repeats, and reading through the entire expansion with 1324 repeats (7941 nucleotides). Using the No-Amp targeted approach, we attained >800x coverage and were able to identify the unexpanded allele, closely estimate expansion size, and assess nucleotide content in a single experiment. We estimate the individuals repeat region was >99% G4C2 content, though we cannot rule out small interruptions.\n\nConclusions: Our findings indicate that long-read sequencing is well suited to characterizing known repeat expansions, and for discovering new disease-causing, disease-modifying, or risk-modifying repeat expansions that have gone undetected with conventional short-read sequencing. The PacBio No-Amp targeted approach may have future potential in clinical and genetic counseling environments. Larger and deeper long-read sequencing studies in C9orf72 expansion carriers will be important to determine heterogeneity and whether the repeats are interrupted by non-G4C2 content, potentially mitigating or modifying disease course or age of onset, as interruptions are known to do in other repeat-expansion disorders. These results have broad implications across all diseases where the genetic etiology remains unclear.

genetics