Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Reconciling disparate estimates of viral genetic diversity during human influenza infections

Deep sequencing can measure viral genetic diversity within human influenza infections, but published studies disagree in their estimates of how much genetic diversity is typically present. One large-scale deep-sequencing study of human influenza reported high levels of shared viral genetic diversity among infected individuals in Hong Kong, but subsequent studies of other cohorts have reported little shared viral diversity. We re-analyze sequencing data from four studies of within-host genetic diversity encompassing more than 500 acute human influenza infections. We identify an anomaly in the Hong Kong data that provides a technical explanation for these discrepancies: read pairs from this study are often split between different biological samples, indicating that some reads are incorrectly assigned. These technical abnormalities explain the high levels of within-host variation and loose transmission bottlenecks reported by this study. Studies without these anomalies consistently report low levels of genetic diversity in acute human influenza infections.

evolutionary biology

The genetics of the mood disorder spectrum: genome-wide association analyses of over 185,000 cases and 439,000 controls

BackgroundMood disorders (including major depressive disorder and bipolar disorder) affect 10-20% of the population. They range from brief, mild episodes to severe, incapacitating conditions that markedly impact lives. Despite their diagnostic distinction, multiple approaches have shown considerable sharing of risk factors across the mood disorders.\n\nMethodsTo clarify their shared molecular genetic basis, and to highlight disorder-specific associations, we meta-analysed data from the latest Psychiatric Genomics Consortium (PGC) genome-wide association studies of major depression (including data from 23andMe) and bipolar disorder, and an additional major depressive disorder cohort from UK Biobank (total: 185,285 cases, 439,741 controls; non-overlapping N = 609,424).\n\nResultsSeventy-three loci reached genome-wide significance in the meta-analysis, including 15 that are novel for mood disorders. More genome-wide significant loci from the PGC analysis of major depression than bipolar disorder reached genome-wide significance. Genetic correlations revealed that type 2 bipolar disorder correlates strongly with recurrent and single episode major depressive disorder. Systems biology analyses highlight both similarities and differences between the mood disorders, particularly in the mouse brain cell types implicated by the expression patterns of associated genes. The mood disorders also differ in their genetic correlation with educational attainment - positive in bipolar disorder but negative in major depressive disorder.\n\nConclusionsThe mood disorders share several genetic associations, and can be combined effectively to increase variant discovery. However, we demonstrate several differences between these disorders. Analysing subtypes of major depressive disorder and bipolar disorder provides evidence for a genetic mood disorders spectrum.

genomics

Time-resolved mapping of genetic interactions to model rewiring of signaling pathways

Context-dependent changes in genetic vulnerabilities are important to understand the wiring of cellular pathways and variations in different environmental conditions. However, methodological frameworks to investigate the plasticity of genetic networks over time or in response to external stresses are lacking. To analyze the plasticity of genetic interactions, we performed an arrayed combinatorial RNAi screen in Drosophila cells at multiple time points and after pharmacological inhibition of Ras signaling activity. Using an image-based morphology assay to capture a broad range of phenotypes, we assessed the effect of 12768 pairwise RNAi perturbations in six different conditions. We found that genetic interactions form in different trajectories and developed an algorithm, termed MODIFI, to analyze how genetic interactions rewire over time. Using this framework, we identified more statistically significant interactions compared to endpoints assays and further observed several examples of context-dependent crosstalk between signaling pathways such as an interaction between Ras and Rel which is dependent on MEK activity.

genomics

Allele-specific binding of RNA-binding proteins reveals functional genetic variants in the RNA

Allele-specific protein-RNA binding is an essential aspect that may reveal functional genetic variants influencing RNA processing and gene expression phenotypes. Recently, genome-wide detection of in vivo binding sites of RNA binding proteins (RBPs) is greatly facilitated by the enhanced UV crosslinking and immunoprecipitation (eCLIP) protocol. Hundreds of eCLIP-Seq data sets were generated from HepG2 and K562 cells during the ENCODE3 phase. These data afford a valuable opportunity to examine allele-specific binding (ASB) of RBPs. To this end, we developed a new computational algorithm, called BEAPR (Binding Estimation of Allele-specific Protein-RNA interaction). In identifying statistically significant ASB sites, BEAPR takes into account UV cross-linking induced sequence propensity and technical variations between replicated experiments. Using simulated data and actual eCLIP-Seq data, we show that BEAPR largely outperforms often-used methods Chi-Squared test and Fishers Exact test. Importantly, BEAPR overcomes the inherent over-dispersion problem of the other methods. Complemented by experimental validations, we demonstrate that ASB events are significantly associated with genetic regulation of splicing and mRNA abundance, supporting the usage of this method to pinpoint functional genetic variants in post-transcriptional gene regulation. Many variants with ASB patterns of RBPs were found as genetic variants with cancer or other disease relevance. About 38% of ASB variants were in linkage disequilibrium with single nucleotide polymorphisms from genome-wide association studies. Overall, our results suggest that BEAPR is an effective method to reveal ASB patterns in eCLIP and can inform functional interpretation of disease-related genetic variants.

bioinformatics

Quantifying the fitness of tissue-specific evolutionary trajectories by harnessing cancer’s repeatability at the genetic level

Cancer is a potentially lethal disease, in which patients with nearly identical genetic backgrounds can develop a similar pathology through distinct combinations of genetic alterations. We aimed to reconstruct the evolutionary process underlying tumour initiation, using the combination of convergence and discrepancies observed across 2,742 cancer genomes from 9 tumour types. We developed a framework using the repeatability of cancer development to score the local malignant adaptation (LMA) of genetic clones, as their potential to malignantly progress and invade their environment of origin. Using this framework, we found that pre-malignant skin and colorectal lesions appeared specifically adapted to their local environment, yet insufficiently for full cancerous transformation. We found that metastatic clones were more adapted to the site of origin than to the invaded tissue, suggesting that genetics may be more important for local progression than for the invasion of distant organs. In addition, we used network analyses to investigate evolutionary properties at the system-level, highlighting that different dynamics of malignant progression can be modelled by such a framework in tumour-type-specific fashion. We find that occurrence-based methods can be used to specifically recapitulate the process of cancer initiation and progression, as well as to evaluate the adaptation of genetic clones to given environments. The repeatability observed in the evolution of most tumour types could therefore be harnessed to better predict the trajectories likely to be taken by tumours and pre-neoplastic lesions in the future.

cancer biology

The effects of demography and genetics on the neutral distribution of quantitative traits

1Neutral models for quantitative trait evolution are useful for identifying phenotypes under selection in natural populations. Models of quantitative traits often assume phenotypes are normally distributed. This assumption may be violated when a trait is affected by relatively few genetic variants or when the effects of those variants arise from skewed or heavy-tailed distributions. Traits such as gene expression levels and other molecular phenotypes may have these properties. To accommodate deviations from normality, models making fewer assumptions about the underlying trait genetics and patterns of genetic variation are needed. Here, we develop a general neutral model for quantitative trait variation using a coalescent approach by extending the framework developed by SO_SCPLOWCHRAIBERC_SCPLOW and LO_SCPLOWANDISC_SCPLOW (2015). This model allows interpretation of trait distributions in terms of familiar population genetic parameters because it is based on the coalescent. We show how the normal distribution resulting from the infinitesimal limit, where the number of loci grows large as the effect size per mutation becomes small, depends only on expected pairwise coalescent times. We then demonstrate how deviations from normality depend on demography through the distribution of coalescence times as well as through genetic parameters. In particular, population growth events exacerbate deviations while bottlenecks reduce them. This model also has practical applications, which we demonstrate by designing an approach to simulate from the null distribution of QST, the ratio of the trait variance between subpopulations to that in the overall population. We further show that it is likely impossible to distinguish sparsity from skewed or heavy-tailed distributions of mutational effects using only trait values sampled from a population. The model analyzed here greatly expands the parameter space for which neutral trait models can be designed.

evolutionary biology

Comparative whole-genome analysis reveals genetic adaptation of the invasive pinewood nematode

Genetic adaptation to new environments is essential for invasive species. To explore the genetic underpinnings of invasiveness of a dangerous invasive species, the pinewood nematode (PWN) Bursaphelenchus xylophilus, we analysed the genome-wide variations of a large cohort of 55 strains isolated from both the native and introduced regions. Comparative analysis showed abundant genetic diversity existing in the nematode, especially in the native populations. Phylogenetic relationships and principal component analysis indicate a dominant invasive population/group (DIG) existing in China and expansion beyond, with few genomic variations. Putative origin and migration paths at a global scale were traced by targeted analysis of rDNA sequences. A progressive loss of genetic diversity was observed along spread routes. We focused on variations with a low frequency allele (<50%) in the native USA population but fixation in DIG, and a total of 25,992 single nuclear polymorphisms (SNPs) were screened out. We found that a clear majority of these fixation alleles originated from standing variation. Functional annotation of these SNP-harboured genes showed that adaptation-related genes are abundant, such as genes that encode for chemoreceptors, proteases, detoxification enzymes, and proteins involved in signal transduction and in response to stresses and stimuli. Some genes under positive selection were predicted. Our results suggest that adaptability to new environments plays essentially roles in PWN invasiveness. Genetic drift, mutation and strong selection drive the nematode to rapidly evolve in adaptation to new environments, which including local pine hosts, vector beetles, commensal microflora and other new environmental factors, during invasion process.

genomics

ForestQC: quality control on genetic variants from next-generation sequencing data using random forest

Next-generation sequencing technology (NGS) enables discovery of nearly all genetic variants present in a genome. A subset of these variants, however, may have poor sequencing quality due to limitations in sequencing technology or in variant calling algorithms. In genetic studies that analyze a large number of sequenced individuals, it is critical to detect and remove those variants with poor quality as they may cause spurious findings. In this paper, we present a statistical approach for performing quality control on variants identified from NGS data by combining a traditional filtering approach and a machine learning approach. Our method uses information on sequencing quality such as sequencing depth, genotyping quality, and GC contents to predict whether a certain variant is likely to contain errors. To evaluate our method, we applied it to two whole-genome sequencing datasets where one dataset consists of related individuals from families while the other consists of unrelated individuals. Results indicate that our method outperforms widely used methods for performing quality control on variants such as VQSR of GATK by considerably improving the quality of variants to be included in the analysis. Our approach is also very efficient, and hence can be applied to large sequencing datasets. We conclude that combining a machine learning algorithm trained with sequencing quality information and the filtering approach is an effective approach to perform quality control on genetic variants from sequencing data.\n\nAuthor SummaryGenetic disorders can be caused by many types of genetic mutations, including common and rare single nucleotide variants, structural variants, insertions and deletions. Nowadays, next generation sequencing (NGS) technology allows us to identify various genetic variants that are associated with diseases. However, variants detected by NGS might have poor sequencing quality due to biases and errors in sequencing technologies and analysis tools. Therefore, it is critical to remove variants with low quality, which could cause spurious findings in follow-up analyses. Previously, people applied either hard filters or machine learning models for variant quality control (QC), which failed to filter out those variants accurately. Here, we developed a statistical tool, ForestQC, for variant QC by combining a filtering approach and a machine learning approach. We applied ForestQC to one family-based whole genome sequencing (WGS) dataset and one general case-control WGS dataset, to evaluate our method. Results show that ForestQC outperforms widely used methods for variant QC by considerably improving the quality of variants. Also, ForestQC is very efficient and scalable to large-scale sequencing datasets. Our study indicates that combining filtering approaches and machine learning approaches enables effective variant QC.

bioinformatics

Genetic population variation and phylogeny of Sinomenium acutum (Menispermaceae) in subtropical China through chloroplast marker

Sinomenium acutum (Menispermaceae) is a traditional Chinese medicine. In recent years, extensive harvesting for medicinal purposes has resulted in a sharp decline in its population. Genetic information is crucial for the proper exploitation and conservation of Sinomenium acutum, but little is known about it at present. In this study, we analyzed 77 samples from 4 populations using four non-coding regions (atpI-atpH, trnQ-5rps16, trnH-psbA, and trnL-trnF) of chloroplast DNA and 14 haplotypes (from C1 to C14) were identified. C1 and C3 were common haplotypes, which were shared by all populations, and C3 was an ancestral haplotype, the rest were rare haplotypes. Obvious phylogeographic structure was not existed inferred by GST / NST test. Mismatch distribution, Tajimas D and Fus FS tests failed to support a rapid demographic expansion in Sinomenium acutum. AMOVA highlighted that the high level of genetic differentiation within population. Low genetic variation among populations illustrated gene flow was not restricted. Genetic diversity analyses demonstrated that the populations of Xuefeng, Dalou, and Daba Mountains were possible refugia localities of Sinomenium acutum. Based on this study, we proposed a preliminary protection strategy for it that C1, C3, C11 and C12 must be collected. These results offer an valuable and useful information for this species of population genetic study as well as further conservation.

molecular biology

A Population Genetic Signature of Polygenic Local Adaptation

Adaptation in response to selection on polygenic phenotypes occurs via subtle allele frequencies shifts at many loci. Current population genomic techniques are not well posed to identify such signals. In the past decade, detailed knowledge about the specific loci underlying polygenic traits has begun to emerge from genome-wide association studies (GWAS). Here we combine this knowledge from GWAS with robust population genetic modeling to identify traits that have undergone local adaptation. Using GWAS data, we estimate the mean additive genetic value for a give phenotype across many populations as simple weighted sums of allele frequencies. We model the expected differentiation of GWAS loci among populations under neutrality to develop simple tests of selection across an arbitrary number of populations with arbitrary population structure. To find support for the role of specific environmental variables in local adaptation we test for correlations with the estimated genetic values. We also develop a general test of local adaptation to identify overdispersion of the estimated genetic values values among populations. This test is a natural generalization of QST /FST comparisons based on GWAS predictions. Finally we lay out a framework to identify the individual populations or groups of populations that contribute to the signal of overdispersion. These tests have considerably greater power than their single locus equivalents due to the fact that they look for positive covariance between like effect alleles. We apply our tests to the human genome diversity panel dataset using GWAS data for six different traits. This analysis uncovers a number of putative signals of local adaptation, and we discuss the biological interpretation and caveats of these results.

Genetics

Genetic evidence for an origin of the Armenians from Bronze Age mixing of multiple populations

The Armenians are a culturally isolated population who historically inhabited a region in the Near East bounded by the Mediterranean and Black seas and the Caucasus, but remain underrepresented in genetic studies and have a complex history including a major geographic displacement during World War One. Here, we analyse genome-wide variation in 173 Armenians and compare them to 78 other worldwide populations. We find that Armenians form a distinctive cluster linking the Near East, Europe, and the Caucasus. We show that Armenian diversity can be explained by several mixtures of Eurasian populations that occurred between [~]3,000 and [~]2,000 BCE, a period characterized by major population migrations after the domestication of the horse, appearance of chariots, and the rise of advanced civilizations in the Near East. However, genetic signals of population mixture cease after [~]1,200 BCE when Bronze Age civilizations in the Eastern Mediterranean world suddenly and violently collapsed. Armenians have since remained isolated and genetic structure within the population developed [~]500 years ago when Armenia was divided between the Ottomans and the Safavid Empire in Iran. Finally, we show that Armenians have higher genetic affinity to Neolithic Europeans than other present-day Near Easterners, and that 29% of the Armenian ancestry may originate from an ancestral population best represented by Neolithic Europeans.

Genetics

Genetic Basis of Transcriptome Diversity in Drosophila melanogaster

Understanding how DNA sequence variation is translated into variation for complex phenotypes has remained elusive, but is essential for predicting adaptive evolution, selecting agriculturally important animals and crops, and personalized medicine. Here, we quantified genome-wide variation in gene expression in the sequenced inbred lines of the Drosophila melanogaster Genetic Reference Panel (DGRP). We found that a substantial fraction of the Drosophila transcriptome is genetically variable and organized into modules of genetically correlated transcripts, which provide functional context for newly identified transcribed regions. We identified regulatory variants for the mean and variance of gene expression, the latter of which could often be explained by an epistatic model. Expression quantitative trait loci for the mean, but not the variance, of gene expression were concentrated near genes. This comprehensive characterization of population scale diversity of transcriptomes and its genetic basis in the DGRP is critically important for a systems understanding of quantitative trait variation.

Genetics

Predicting genetic interactions from Boolean models of biological networks

Genetic interaction can be defined as a deviation of the phenotypic quantitative effect of a double gene mutation from the effect predicted from single mutations using a simple (e.g., multiplicative or linear additive) statistical model. Experimentally characterized genetic interaction networks in model organisms provide important insights into relationships between different biological functions. We describe a computational methodology allowing to systematically and quantitatively characterize a Boolean mathematical model of a biological network in terms of genetic interactions between all loss of function and gain of function mutations with respect to all model phenotypes or outputs. We use the probabilistic framework defined in MaBoSS software, based on continuous time Markov chains and stochastic simulations. In addition, we suggest several computational tools for studying the distribution of double mutants in the space of model phenotype probabilities. We demonstrate this methodology on three published models for each of which we derive the genetic interaction networks and analyze their properties. We classify the obtained interactions according to their class of epistasis, dependence on the chosen initial conditions and phenotype. The use of this methodology for validating mathematical models from experimental data and designing new experiments is discussed.{dagger}

Genetics

Genetic basis underlying connection between hyperglycemia and dyslipidemia in Apoe-deficient mice

Individuals with dyslipidemia often develop type 2 diabetes, and diabetic patients often have dyslipidemia. It remains to be determined whether there are genetic connections between the 2 disorders. A female F2 cohort, generated from BALB/cJ (BALB) and SM/J (SM) Apoe-deficient (Apoe-/-) strains, was fed a Western diet for 12 weeks. Fasting plasma glucose and lipid levels were measured before and after Western diet feeding. 144 genetic markers across the entire genome were used for analysis. One significant QTL on chromosome 9, named Bglu17 [26.4 cM, logarithm of odds ratio (LOD): 5.4], and 3 suggestive QTLs were identified for fasting glucose levels. The suggestive QTL near the proximal end of chromosome 9 (2.4 cM, LOD: 3.12) was detected when mice were fed chow or Western diet and named Bglu16. Bglu17 coincided with a significant QTL for HDL and a suggestive QTL for non-HDL cholesterol levels. Plasma glucose levels were inversely correlated with HDL but positively correlated with non-HDL cholesterol levels in F2 mice fed either diet. A significant correlation between fasting glucose and triglyceride levels was observed on the Western but not chow diet. Haplotype analysis revealed that \"lipid genes\" Sik3 and Apoc3 were probable candidates for Bglu17. We have identified multiple QTLs for fasting glucose and lipid levels. The colocalization of QTLs for both phenotypes and the sharing of potential causal genes suggest that dyslipidemia and type 2 diabetes are genetically connected.\n\nArticle SummaryPatients with dyslipidemia often develop type 2 diabetes, and diabetic patients often have dyslipidemia. It remains unknown whether there are genetic connections between the 2 disorders. Using a female F2 cohort derived from BALB/cJ and SM/J Apoe-deficient mice, we identified one significant QTL on chromosome 9, named Bglu17, and 3 suggestive QTLs were identified for fasting glucose levels. Bglu17 coincided with a significant QTL for HDL and a suggestive QTL for non-HDL levels. Plasma glucose levels were significantly correlated with HDL and non-HDL levels in F2 mice. Haplotype analysis revealed Sik3 and Apoc3 were probable candidates for both QTLs.

Genetics

A method to estimate the contribution of regional genetic associations to complex traits from summary association statistics

Despite considerable efforts, known genetic associations only explain a small fraction of predicted heritability. Regional associations combine information from multiple contiguous genetic variants and can improve variance explained at established association loci. However, regional associations are not easily amenable to estimation using summary association statistics because of sensitivity to linkage disequilibrium (LD). We now propose a novel method to estimate phenotypic variance explained by regional associations using summary statistics while accounting for LD. Our method is asymptotically equivalent to multiple regression models when no interaction or haplotype effects are present. It has multiple applications, such as ranking of genetic regions according to variance explained or comparison of variance explained by to or more regions. Using height and BMI data from the Health Retirement Study (N=7,776), we show that most genetic variance lies in a small proportion of the genome and that previously identified linkage peaks have higher than expected regional variance.

Genetics

Molecular genetic contributions to social deprivation and household income in UK Biobank (n = 112,151)

Individuals with lower socio-economic status (SES) are at increased risk of physical and mental illnesses and tend to die at an earlier age [1-3]. Explanations for the association between SES and health typically focus on factors that are environmental in origin [4]. However, common single nucleotide polymorphisms (SNPs) have been found collectively to explain around 18% (SE = 5%) of the phenotypic variance of an area-based social deprivation measure of SES [5]. Molecular genetic studies have also shown that physical and psychiatric diseases are at least partly heritable [6]. It is possible, therefore, that phenotypic associations between SES and health arise partly due to a shared genetic etiology. We conducted a genome-wide association study (GWAS) on social deprivation and on household income using the 112,151 participants of UK Biobank. We find that common SNPs explain 21% (SE = 0.5%) of the variation in social deprivation and 11% (SE = 0.7%) in household income. Two independent SNPs attained genome-wide significance for household income, rs187848990 on chromosome 2, and rs8100891 on chromosome 19. Genes in the regions of these SNPs have been associated with intellectual disabilities, schizophrenia, and synaptic plasticity. Extensive genetic correlations were found between both measures of socioeconomic status and illnesses, anthropometric variables, psychiatric disorders, and cognitive ability. These findings show that some SNPs associated with SES are involved in the brain and central nervous system. The genetic associations with SES are probably mediated via other partly-heritable variables, including cognitive ability, education, personality, and health.

Genetics

Novel Genetic Risk factors for Asthma in African American Children: Precision Medicine and The SAGE II Study.

BackgroundAsthma, an inflammatory disorder of the airways, is the most common chronic disease of children worldwide. There are significant racial/ethnic disparities in asthma prevalence, morbidity and mortality among U.S. children. This trend is mirrored in obesity, which may share genetic and environmental risk factors with asthma. The majority of asthma biomedical research has been performed in populations of European decent.\n\nObjectiveWe sought to identify genetic risk factors for asthma in African American children. We also assessed the generalizability of genetic variants associated with asthma in European and Asian populations to African American children.\n\nMethodsOur study population consisted of 1227 (812 asthma cases, 415 controls) African American children with genome-wide single nucleotide polymorphism (SNP) data. Logistic regression was used to identify associations between SNP genotype and asthma status.\n\nResultsWe identified a novel variant in the PTCHD3 gene that is significantly associated with asthma (rs660498, p = 2.2 x10-7) independent of obesity status. Fewer than 5% of previously reported asthma genetic associations identified in European populations replicated in African Americans.\n\nConclusionsOur identification of novel variants associated with asthma in African American children, coupled with our inability to replicate the majority of findings reported in European Americans, underscores the necessity for including diverse populations in biomedical studies of asthma.

Genetics

A method to exploit the structure of genetic ancestry space to enhance case-control studies

One goal of human genetics is to understand the genetic basis of disease, a challenge for diseases of complex inheritance because risk alleles are few relative to the vast set of benign variants. Risk variants are often sought by association studies in which allele frequencies in cases are contrasted with those from population-based samples used as controls. In an ideal world we would know population-level allele frequencies, releasing researchers to focus on case subjects. We argue this ideal is possible, at least theoretically, and we outline a path to achieving it in reality. If such a resource were to exist, it would yield ample savings and would facilitate the effective use of data repositories by removing administrative and technical barriers. We call this concept the Universal Control Repository Network (UNICORN), a means to perform association analyses without necessitating direct access to individual-level control data. Our approach to UNICORN uses existing genetic resources and various statistical tools to analyze these data, including hierarchical clustering with spectral analysis of ancestry; and empirical Bayesian analysis along with Gaussian spatial processes to estimate ancestry-specific allele frequencies. We demonstrate our approach using tens of thousands of controls from studies of Crohns disease, showing how it controls false positives, provides power similar to that achieved when all control data are directly accessible, and enhances power when control data are limiting or even imperfectly matched ancestrally. These results highlight how UNICORN can enable reliable, powerful and convenient genetic association analyses without access to the individual level data.

Genomics