Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

Seasonal diets overwhelm host species in shaping the gut microbiota of Yak and Tibetan sheep

Host genetics and environmental factors can both shaping composition of gut microbiota, yet which factors are more important is still under debating. Yak (Bos grunniens) and Tibetan sheep (Ovis aries) are very different from the size and genetics. Nomadic Tibetan people keep them as main livestock and feeding them with same grazing systems, which provide a good opportunity to study the effects of diet and host species on gut microbiome. We collected fecal samples from yaks and Tibetan sheeps at different seasons when they were feed with different diets. Illumina data showed that major bacterial phyla of both animals are Bacteroidetes and Firmicutes, which agree with the previous reports. And the season effect had a higher impact on the gut microbiota than that of host species, though the animals are taxonomically distinguished each other at subfamily level. Since that the animal grazing differently at different seasons, this study indicated that diet can trump the host genetics even at higher taxonomic level. This finding provides a cautionary note for the researchers to link host genetics to the composition and function of the gut microbiota.\n\nImportanceYak and Tibetan sheep are very different from the size and genetics (from different sub-family). Nomadic Tibetan people keep them as main livestock and feeding them with same grazing systems, which provide a good opportunity to study the effects of diet and host species on the gut microbiota. Results indicated that diet can trump the host genetics even at higher taxonomic level. This finding provides a cautionary note for the researchers to link host genetics to the composition and function of the gut microbiota.

ecology

A multivariate genome-wide association analysis of 10 LDL subfractions, and their response to statin treatment, in 1868 Caucasians

We conducted a genome-wide association analysis of 7 subfractions of low density lipoproteins (LDLs) and 3 subfractions of intermediate density lipoproteins (IDLs) measured by gradient gel electrophoresis, and their response to statin treatment, in 1868 individuals of European ancestry from the Pharmacogenomics and Risk of Cardiovascular Disease study. Our analyses identified four previously-implicated loci (SORT1, APOE, LPA, and CETP) as containing variants that are very strongly associated with lipoprotein subfractions (log10Bayes Factor > 15). Subsequent conditional analyses suggest that three of these (APOE, LPA and CETP) likely harbor multiple independently associated SNPs. Further, while different variants typically showed different characteristic patterns of association with combinations of subfractions, the two SNPs in CETP show strikingly similar patterns - both in our original data and in a replication cohort - consistent with a common underlying molecular mechanism. Notably, the CETP variants are very strongly associated with LDL subfractions, despite showing no association with total LDLs in our study, illustrating the potential value of the more detailed phenotypic measurements. In contrast with these strong subfraction associations, genetic association analysis of subfraction response to statins showed much weaker signals (none exceeding log10 Bayes Factor of 6). However, two SNPs (in APOE and LPA) previously-reported to be associated with LDL statin response do show some modest evidence for association in our data, and the subfraction response profiles at the LPA SNP are consistent with the LPA association, with response likely being due primarily to resistance of Lp(a) particles to statin therapy. An additional important feature of our analysis is that, unlike most previous analyses of multiple related phenotypes, we analyzed the subfractions jointly, rather than one at a time. Comparisons of our multivariate analyses with standard univariate analyses demonstrate that multivariate analyses can substantially increase power to detect associations. Software implementing our multivariate analysis methods is available at http://stephenslab.uchicago.edu/software.html\n\nAuthor SummaryLevels of plasma lipids and lipoproteins are related to risk of cardiovascular disease (CVD), and because of this, considerable attention has been devoted to genetic association analyses of lipid-related measures. In addition, motivated by the fact that statins are widely prescribed to lower plasma low density lipoprotein (LDL) cholesterol and CVD risk, and that response to statins has a genetic component, several studies have searched for genetic associations with response of lipid related phenotypes to statin treatment. Here, in 1868 individuals of European ancestry from the Pharmacogenomics and Risk of Cardiovascular Disease study, we have conducted genetic association analyses of 7 subfractions of LDLs and 3 subfractions of intermediate density lipoproteins (IDLs) measured by gradient gel electrophoresis, and their response to statin treatment. These phenotypic measurements offer higher resolution information on LDLs and IDLs than available previously. Therefore, our study provides a more detailed picture of association with the entire IDL/LDL subfraction profile than any prior genetic association studies of either lipid-related measures or their response to statin treatment. Moreover, unlike most previous analyses of multiple related measurements, we analyzed the subfractions jointly, rather than one at a time. Our results demonstrate that joint analyses of related measurements can considerably increase power to detect associations compared with conventional univariate analyses.

Genetics

Mutation rate estimation for 15 autosomal STR loci in a large population from Mainland China

STR, short trandem repeats, is well known as a type of powerful genetic marker and widely used in studying human population genetics. Compared with the conventional genetic markers, the mutation rate of STR is higher. Additionally, the mutations of STR loci do not lead to genetic inconsistencies between the genotypes of parents and children; therefore, the analysis of STR mutation is more suited to assess the population mutation. In this study, we focused on 15 autosomal STR loci (D8S1179, D21S11, D7S820, CSF1PO, D3S1358, TH01, D13S317, D16S539, D2S1338, D19S433, vWA, TPOX, D18S51, D5S818, FGA). DNA samples from a total of 42416 unrelated healthy individuals (19037 trios) from the population of Mainland China collected between Jan 2012 and May 2014 were successfully investigated. In our study, the allele frequencies, paternal mutation rates, maternal mutation rates and average mutation rates were detected in the 15 STR loci. Furthermore, we also investigated the relationship between paternal ages, maternal ages, pregnant time, area and average mutation rate. We found that paternal mutation rate is higher than maternal mutation rate and the paternal, maternal, and average mutation rates have a positive correlation with paternal ages, maternal ages and times respectively. Additionally, the average mutation rates of coastal areas are higher than that of inland areas. Overall, these results suggest that the 15 autosomal STR loci can provide highly informative polymorphic data for population genetic assessment in Mainland China, as well as confirm and extend the application of STR analysis in population genetics.

Genetics

Mapping and Inheritance analysis of a novel dominant rice male sterility mutant, OsDMS-1

We found a rice dominant genetic male sterile mutant OsDMS-1 from the tissue culture regenerated offspring of Zhonghua 11 (japonica rice). Compared to wild Zhonghua 11, OsDMS-1 mutant anthers were thinner and whiter, and could not release any pollen although the glume opened normally; most of the mutant pollen was small and malformed, and could not be stained by iodine treatment; a paraffin section assay showed the degradation of OsDMS-1 mutant tapetum was delayed, with no accumulation of starch in the mutant pollen, ultimately leading to pollen abortion. Classical genetic analysis indicated that only one dominant gene was controlling the sterility in the OsDMS-1 mutant. However, molecular mapping suggested three loci simultaneously control male sterility in this mutant: OsDMS-1A, flanked by InDel markers C1D4 and C1D5 with a genetic distance of 0.15 and 0.30 cM, respectively; OsDMS-1B, flanked by InDel markers C2D3 and C2D10 with a genetic distance of 0.44 and 0.88 cM, respectively; OsDMS-1C, flanked by InDel markers 0315 and C3D3 with a genetic distance of 0.44 and 0.88 cM, respectively. Molecular mapping disagreed with classical genetic analysis about the number of controlling genes in the OsDMS-1 mutant, indicating a novel mechanism underlying sterility in OsDMS-1. We present two hypotheses to explain this novel inheritance behavior: one is described as Parent-Originated Loci Tying Inheritance (POLTI); or the hypothesis is described as Loci Recombination Lethal (LRL).\n\nKey messageThree loci, which were localized on the chromosomes 1, 2 and 3 respectively, simultaneously control a dominant rice male sterility in this mutant: OsDMS-1.

Genetics

Extreme distribution of deleterious variation in a historically small and isolated population-insights from the Greenlandic Inuit

The genetic consequences of a severe bottleneck on genetic load in humans are widely disputed. Based on exome sequencing of 18 Greenlandic Inuit we show that the Inuit have undergone a severe ~20,000 yearlong bottleneck. This has led to a markedly more extreme distribution of deleterious alleles than seen for any other human population. Compared to populations with much larger population sizes, we see an overall reduction in the number of variable sites, increased numbers of fixed sites, a lower heterozygosity, and increased mean allele frequency as well as more homozygous deleterious genotypes. This means, that the Inuit population is the perfect population to examine the effect of a bottleneck on genetic load. Compared to the European, Asian and African populations, we do not observe a difference in the overall number of derived alleles. In contrast, using proxies for genetic load we find that selection has acted less efficiently in the Inuit, under a recessive model. This fits with our simulations that predict a similar number of derived alleles but a true higher genetic load for the Inuit regardless of the genetic model. Finally, we find that the Inuit population has a great potential for mapping of disease-causing variants that are rare in large populations. In fact, we show that these alleles are more likely to be common, and thus easy to map, in the Inuit than in the Finnish and Latino populations; populations considered highly valuable for mapping studies due to recent bottleneck events.

Genetics

Why are frameshift homologs widespread within and across species?

Frameshift protein sequences encoded by alternative reading frames of coding genes have been considered meaningless, and frameshift mutations have been considered of little importance for the molecular evolution of coding genes and proteins. However, functional frameshifts have been found widely existing. It was puzzling how a frameshift protein kept its structure and functionality while its amino-acid sequence was changed substantially. Here we show that frame similarities between frameshifts and wild types are higher than random similarities and are defined at the genetic code, gene, and genome levels. In the standard genetic code, frameshift codon substitutions are more conservative than random substitutions. The frameshift tolerability of the standard genetic code ranks in the top 2.0-3.5% of alternative genetic codes, showing that the genetic code is nearly optimal for frameshift tolerance. Furthermore, frameshift-resistant codons (codon pairs) appear more frequently than expected in many genes and certain genomes, showing that the frameshift optimality is reflected not only in the genetic code but more importantly, in its allowance of further optimizing the frameshift tolerance of a particular gene or genome, which shed light on the role of frameshift mutations in molecular and genomic evolution.

Genetics

Pleiotropy-robust Mendelian Randomization

BackgroundThe potential of Mendelian Randomization studies is rapidly expanding due to (i) the growing power of GWAS meta-analyses to detect genetic variants associated with several exposures, and (ii) the increasing availability of these genetic variants in large-scale surveys. However, without a proper biological understanding of the pleiotropic working of genetic variants, a fundamental assumption of Mendelian Randomization (the exclusion restriction) can always be contested.\n\nMethodsWe build upon and synthesize recent advances in the econometric literature on instrumental variables (IV) estimation that test and relax the exclusion restriction. Our Pleiotropy-robust Mendelian Randomization (PRMR) method first estimates the degree of pleiotropy, and in turn corrects for it. If a sample exists for which the genetic variants do not affect the exposure, and pleiotropic effects are homogenous, PRMR obtains unbiased estimates of causal effects in case of pleiotropy.\n\nResultsSimulations show that existing MR methods produce biased estimators for realistic forms of pleiotropy. Under the aforementioned assumptions, PRMR produces unbiased estimators. We illustrate the practical use of PRMR by estimating the causal effect of (i) cigarettes smoked per day on Body Mass Index (BMI); (ii) prostate cancer on self-reported health, and (iii) educational attainment on BMI in the UK Biobank data.\n\nConclusionsPRMR allows for instrumental variables that violate the exclusion restriction due to pleiotropy, and corrects for pleiotropy in the estimation of the causal effect. If the degree of pleiotropy is unknown, PRMR can still be used as a sensitivity analysis.\n\nKey messagesO_LIIf genetic variants have pleiotropic effects, causal estimates of Mendelian Randomization studies will be biased.\nC_LIO_LIPleiotropy-robust Mendelian Randomization (PRMR) produces unbiased causal estimates in case (i) a subsample can be identified for which the genetic variants do not affect the exposure, and (ii) pleiotropic effects are homogenous.\nC_LIO_LIIf such a subsample does not exist, PRMR can still routinely be reported as a sensitivity analysis in any MR analysis.\nC_LIO_LIIf pleiotropic effects are not homogenous, PRMR can be used as an informal test to gauge the exclusion restriction.\nC_LI

Genetics

A global perspective of codon usage

Codon usage in 2730 genomes is analyzed for evolutionary patterns in the usage of synonymous codons and amino acids across prokaryotic and eukaryotic taxa. We group genomes together that have similar amounts of intra-genomic bias in their codon usage, and then compare how usage of particular different codons is diversified across each genome group, and how that usage varies from group to group. Inter-genomic diversity of codon usage increases with intra-genomic usage bias, following a universal pattern. The frequencies of the different codons vary in robust mutual correlation, and the implied synonymous codon and amino acid usages drift together. This kind of correlation indicates that the variation of codon usage across organisms is chiefly a consequence of lateral DNA transfer among diverse organisms. The group of genomes with the greatest intra-genomic bias comprises two distinct subgroups, with each one restricting its codon usage to essentially one unique half of the genetic code table. These organisms include eubacteria and archaea thought to be closest to the hypothesized last universal common ancestor (LUCA). Their codon usages imply genetic diversity near the hypothesized base of the tree of life. There is a continuous evolutionary progression across taxa from the two extremely diversified usages toward balanced usage of different codons (as approached, e.g. in mammals). In that progression, codon frequency variations are correlated as expected from a blending of the two extreme codon usages seen in prokaryotes.\n\nAUTHOR SUMMARYThe redundancy intrinsic to the genetic code allows different amino acids to be encoded by up to six synonymous codons. Genomes of different organisms prefer different synonymous codons, a phenomenon known as codon usage bias. The phenomenon of codon usage bias is of fundamental interest for evolutionary biology, and is important in a variety of applied settings (e.g., transgene expression). The spectrum of codon usage biases seen in current organisms is commonly thought to have arisen by the combined actions of mutations and selective pressures. This view focuses on codon usage in specific genomes and the consequences of that usage for protein expression.\n\nHere we investigate an unresolved question of molecular genetics: are there global rules governing the usage of synonymous codons made by genomic DNA across organisms? To answer this question, we employed a data-driven approach to surveying 2730 species from all kingdoms of the tree of life in order to classify their codon usage. A first major result was that the large majority of these organisms use codons rather uniformly on the genome-wide scale, without giving preference to particular codons among possible synonymous alternatives. A second major result was that two compartments of codon usage seem to co-exist and to be expressed in different proportions by different organisms. As such, we investigate how individual different codons are used in different organisms from all taxa. Whereas codon usage is generally believed to be the evolutionary result of both mutations and natural selection, our results suggest a different perspective: the usage of different codons (and amino acids) by different organisms follows a superposition of two distinct patterns of usage. One distinction locates to the third base pair of all different codons, which in one pattern is U or A, and in the other pattern is G or C. This result has two major implications: (1) the variation of codon usage as seen across different organisms is best accounted for by lateral gene transfer among diverse organisms; (2) the organisms that are by protein homology grouped near the base of the tree of life comprise two genetically distinct lineages.\n\nWe find that, over evolutionary time, codon usages have converged from two distinct, non-overlapping usages (e.g., as evident in bacteria and archaea) to a near-uniform, balanced usage of synonymous codons (e.g., in mammals). This shows that the variations of codon (and amino acid) biases reveal a distinct evolutionary progression. We also find that codon usage in bacteria and archaea is most diverse between organisms thought to be closest to the hypothesized last universal common ancestor (LUCA). The dichotomy in codon (and amino acid usages) present near the origin of the current tree of life might provide information about the evolutionary development of the genetic code.

Genetics

Epistatic networks jointly influence phenotypes related to metabolic disease and gene expression in Diversity Outbred mice

Genetic studies of multidimensional phenotypes can potentially link genetic variation, gene expression, and physiological data to create multi-scale models of complex traits. Multi-parent populations provide a resource for developing methods to understand these relationships. In this study, we simultaneously modeled body composition, serum biomarkers, and liver transcript abundances from 474 Diversity Outbred mice. This population contained both sexes and two dietary cohorts. Using weighted gene co-expression network analysis (WGCNA), we summarized transcript data into functional modules which we then used as summary phenotypes representing enriched biological processes. These module phenotypes were jointly analyzed with body composition and serum biomarkers in a combined analysis of pleiotropy and epistasis (CAPE), which inferred networks of epistatic interactions between quantitative trait loci that affect one or more traits. This network frequently mapped interactions between alleles of different ancestries, providing evidence of both genetic synergy and redundancy between haplotypes. Furthermore, a number of loci interacted with sex and diet to yield sex-specific genetic effects. We were also able to identify alleles that potentially protect individuals from the effects of a high-fat diet. Although the epistatic interactions explained small amounts of trait variance, the combination of directional interactions, allelic specificity, and high genomic resolution provided context to generate hypotheses for the roles of specific genes in complex traits. Our approach moves beyond the cataloging of single loci to infer genetic networks that map genetic etiology by simultaneously modeling all phenotypes.

genetics

Whole genome linkage analysis in a large Brazilian multigenerational family reveals distinct linkage signals for Bipolar Disorder and Depression.

Both common and rare genetic variation play a role in the causes for mood disorders. Very large families pose unique opportunities and analytical challenges but may provide a way to identify regions and mutations associated with mood disorders. We identified a family with a high prevalence (~30%) of mood disorders in a rural village in Brazil, featuring decreasing age of onset over generations. The pattern of inheritance was complex with 32 Bipolar type I cases, 11 Bipolar type II and 59 recurrent and/or severe Depression cases in addition to other phenotypes. We enrolled 333 participants with DNA samples from a broader pedigree of 960 subjects for genotyping using the Affymetrix 10K array. Non-parametric linkage was carried out via MERLIN and parametric with both MERLIN and MCLINKAGE. We exome sequenced a subset of the family (n=27) in order to identify rare variation within the linkage regions shared by affected family members. We identified four genome wide significant and four suggestive linkage regions on chromosomes 1, 2, 3, 11 and 12 for different phenotype definitions. However, no region received strong joint support in both the parametric and non-parametric analyses. Exome sequencing revealed potential deleterious variants in 11p15.4 for MDD and 1q21.1-1q21.3 and 12p23.1-p22.3, implicated in cell signaling, adhesion, translation and neurogenesis processes. Overall, our results suggest promising, but not definitive or confirmed evidence, that rare genetic variation contributes to the high prevalence of mood disorders in this multi-generational family. We note that a substantial role for common genetic variation is likely given the strength of the linkage signals observed.\n\nThe World Health Organisation reports depression and bipolar disorder as the second and seventh most important causes of years lost due to disability worldwide[1]. The heritability of bipolar disorder is between 60-90% with a lower but still substantial heritability for major depression (40-45%) [2]; [3]. First-degree relatives of bipolar disorder probands have a 5-10 fold increase in risk of developing the illness compared to relatives of controls but also show a three fold increase in unipolar depression, indicating that bipolar disorder does not \"breed true\" [4]. Large collaborative genome-wide association studies (GWAS) have uncovered several common genetic variants of small effect [5]. Genomewide estimates of heritability suggest that up to 60% of the genetic risk is contributed by common variants [6]. Overall, the current picture for bipolar disorder (and almost all complex traits) is a genetic architecture formed of both common and rare variants.\n\nLinkage studies have been pursued on the basis that there may be variants of greater effect shared between and within affected families. However these studies have usually focused on collections of comparatively small families or sib pairs and few consistent findings have emerged [7]. Large multigenerational families (e. g. of >30 affected individuals) theoretically offer a powerful means for mapping complex disease loci that are individually rare but common in a single family. These loci may be more highly penetrant and of larger effect than loci found with GWAS [8]. Here we report the results of the Brazilian Bipolar Family (BBF) study on a five-generation family of 639 members of which 333 were enrolled in the current analyses. Our objectives were to perform a linkage analysis with genome coverage and try to identify new genes/mutations related to bipolar and other mood disorders in the family. Here we report our findings and preliminary results of sequencing of linkage regions.

genetics

Mitochondrial Dual-coding Genes in Trypanosoma brucei

Trypanosoma brucei is transmitted between mammalian hosts by the tsetse fly. In the mammal, they are exclusively extracellular, continuously replicating within the bloodstream. During this stage, the mitochondrion lacks a functional electron transport chain (ETC). Successful transition to the fly, requires activation of the ETC and ATP synthesis via oxidative phosphorylation. This life cycle leads to a major problem: in the bloodstream, the mitochondrial genes are not under selection and are subject to genetic drift that endangers their integrity. Exacerbating this, T. brucei undergoes repeated population bottlenecks as they evade the host immune system that would create additional forces of genetic drift. These parasites possess several unique genetic features, including RNA editing of mitochondrial transcripts. RNA editing creates open reading frames by the guided insertion and deletion of U-residues within the mRNA. A major question in the field has been why this metabolically expensive system of RNA editing would evolve and persist. Here, we show that many of the edited mRNAs can alter the choice of start codon and the open reading frame by alternative editing of the 5 end. Analyses of mutational bias indicate that six of the mitochondrial genes may be dual-coding and that RNA editing allows access to both reading frames. We hypothesize that dual-coding genes can protect genetic information by essentially hiding a non-selected gene within one that remains under selection. Thus, the complex RNA editing system found in the mitochondria of trypanosomes provides a unique molecular strategy to combat genetic drift in non-selective conditions.\n\nAuthor SummaryIn African trypanosomes, many of the mitochondrial mRNAs require extensive RNA editing before they can be translated. During this process, each edited transcript can undergo hundreds of cleavage/ligation events as U-residues are inserted or deleted to generate a translatable open reading frame. A major paradox has been why this incredibly metabolically expensive process would evolve and persist. In this work, we show that many of the mitochondrial genes in trypanosomes are dual-coding, utilizing different reading frames to potentially produce two very different proteins. Access to both reading frames is made possible by alternative editing of the 5 end of the transcript. We hypothesize that dual-coding genes may work to protect the mitochondrial genes from mutations during growth in the mammalian host, when many of the mitochondrial genes are not being used. Thus, the complex RNA editing system may be maintained because it provides a unique molecular strategy to combat genetic drift.

genetics

Coalescent theory of migration network motifs

Natural populations display a variety of spatial arrangements, each potentially with a distinctive impact on genetic diversity and genetic differentiation among subpopulations. Although the spatial arrangement of populations can lead to intricate migration networks, theoretical developments have focused mainly on a small subset of such networks, emphasizing the island-migration and stepping-stone models. In this study, we investigate all small network motifs: the set of all possible migration networks among populations subdivided into at most four subpopulations. For each motif, we use coalescent theory to derive expectations for three quantities that describe genetic variation: nucleotide diversity, FST, and half-time to equilibrium diversity. We describe the impact of network properties on these quantities, finding that motifs with a large mean node degree have the largest nucleotide diversity and the longest time to equilibrium, whereas motifs with small density have the largest FST. In addition, we show that the motifs whose pattern of variation is most strongly influenced by loss of a connection or a subpopulation are those that can be split easily into several disconnected components. We illustrate our results using two example datasets--sky island birds of genus Brachypteryx and Indian tigers--identifying disturbance scenarios that produce the greatest reduction in genetic diversity; for tigers, we also compare the benefits of two assisted gene flow scenarios. Our results have consequences for understanding the effect of geography on genetic diversity and for designing strategies to alter population migration networks to maximize genetic variation in the context of conservation of endangered species.

genetics

Genome-wide Analysis of Insomnia (N=1,331,010) Identifies Novel Loci and Functional Pathways

Insomnia is the second-most prevalent mental disorder, with no sufficient treatment available. Despite a substantial role of genetic factors, only a handful of genes have been implicated and insight into the associated neurobiological pathways remains limited. Here, we use an unprecedented large genetic association sample (N=1,331,010) to allow detection of a substantial number of genetic variants and gain insight into biological functions, cell types and tissues involved in insomnia complaints. We identify 202 genome-wide significant loci implicating 956 genes through positional, eQTL and chromatin interaction mapping. We show involvement of the axonal part of neurons, of specific cortical and subcortical tissues, and of two specific cell-types in insomnia: striatal medium spiny neurons and hypothalamic neurons. These cell-types have been implicated previously in the regulation of reward processing, sleep and arousal in animal studies, but have never been genetically linked to insomnia in humans. We found weak genetic correlations with other sleep-related traits, but strong genetic correlations with psychiatric and metabolic traits. Mendelian randomization identified causal effects of insomnia on specific psychiatric and metabolic traits. Our findings reveal key brain areas and cells implicated in the neurobiology of insomnia and its related disorders, and provide novel targets for treatment.

genetics

Proper Conditional Analysis in the Presence of Missing Data Identified Novel Independently Associated Low Frequency Variants in Nicotine Dependence Genes

Meta-analysis of genetic association studies increases sample size and the power for mapping complex traits. Existing methods are mostly developed for datasets without missing values. In practice, genotype imputation is not always effective, e.g. when targeted genotyping/sequencing assays are used or when the un-typed genetic variant is rare. Therefore, contributed summary statistics often contain missing values. Naive extensions of existing methods either replace missing summary statistics with 0 or discard studies with missing data. These approaches can bias genetic effect estimates and lead to seriously inflated type-I or II errors in conditional analysis, which is a critical tool for identifying independently associated variants.\n\nTo address this challenge and complement imputation methods, we developed a method to combine summary statistics across participating studies and consistently estimate joint effects, even when the contributed summary statistics contain large amount of missing values. Based on this estimator, we propose a score statistic we call PCBS (partial correlation based score statistic) for conditional analysis of single-variant and gene-level associations. Through extensive analysis of simulated and real data, we showed that the new method produces well-calibrated type-I errors and is substantially more powerful than existing approaches. We applied the proposed approach to analyze the CHRNA5-CHRNB4-CHRNA3 locus in a large-scale meta-analysis for cigarettes-per-day. Using the new method, we identified three novel variants, independent of known association signals, which were otherwise missed by alternative methods. Together, the phenotypic variance explained by these variants is .46%, improving that of previously reported associations by 17%. These findings illustrate the extent of locus allelic heterogeneity and can help pinpoint causal variants.\n\nAUTHOR SUMMARYIt is of great interest to estimate the joint and conditional effects of multiple correlated variants from large scale meta-analysis, in order to fine map causal variants and understand the genetic architecture for complex traits. The contributed summary statistics from participating studies in a meta-analysis often contain missing values, as the imputation methods are not often effective, especially when the underlying genetic variant is rare or the participating studies use targeted genotyping array that is not suitable for imputation. Existing meta-analysis methods do not properly handle missing data, and can incorrectly estimate correlations between score statistics. As a result, they can produce highly biased estimates of joint effects and highly inflated type-I errors for conditional analysis, which will in turn result in overestimated phenotypic variance explained and incorrect identification of causal variants. We systematically evaluated this bias and proposed a novel partial correlation based score statistic. The new statistic has valid type-I errors for conditional analysis and much higher power than the existing methods, even when the contributed summary statistics in the meta-analysis contain a large fraction of missing values. We expect this method to be highly useful in the sequencing age for complex trait genetics.

genetics

The asexual genome of Drosophila

The rate of recombination affects the mode of molecular evolution. In high-recombining sequence, the targets of selection are individual genetic loci; under low recombination, selection collectively acts on large, genetically linked genomic segments. Selection under linkage can induce clonal interference, a specific mode of evolution by competition of genetic clades within a population. This mode is well known in asexually evolving microbes, but has not been traced systematically in an obligate sexual organism. Here we show that the Drosophila genome is partitioned into two modes of evolution: a local interference regime with limited effects of genetic linkage, and an interference condensate with clonal competition. We map these modes by differences in mutation frequency spectra, and we show that the transition between them occurs at a threshold recombination rate that is predictable from genomic summary statistics. We find the interference condensate in segments of low-recombining sequence that are located primarily in chromosomal regions flanking the centromeres and cover about 20% of the Drosophila genome. Condensate regions have characteristics of asexual evolution that impact gene function: the efficacy of selection and the speed of evolution are lower and the genetic load is higher than in regions of local interference. Our results suggest that multicellular eukaryotes can harbor heterogeneous modes and tempi of evolution within one genome. We argue that this variation generates selection on genome architecture.\n\nAuthor SummaryThe Drosophila genome is an ideal system to study how the rate of recombination affects molecular evolution. It harbors a wide range of local recombination rates, and its high-recombining parts show broad signatures of adaptive evolution. The low-recombining parts, however, have remained dark genomic matter that has been omitted from most studies on the inference of selection. Here we show that these genomic regions evolve in a different way, which involves clonal competition and is akin to the evolution of asexual systems. This regime shows a lower efficacy of selection, a lower speed of evolution, and a higher genetic load than high-recombining regions. We argue these evolutionary differences have functional consequences: protein stability and protein expression are gene traits likely to be partially compromised by low recombination rates.

genetics

Factor Structure and Heritability of Obsessive-Compulsive Traits in Children and Adolescents in the General Population

BackgroundObsessive-compulsive disorder (OCD) is a heritable childhood-onset psychiatric disorder that may represent the extreme of obsessive-compulsive (OC) traits that are widespread in the general population. We studied the factor structure and heritability of the Toronto Obsessive Compulsive Scale (TOCS), a new measure designed to assess traits associated with OCD in children and adolescents. We also examined the degree to which genetic effects are unique and shared between dimensions.\n\nMethodsOC traits were measured using the TOCS in 16,718 children and adolescents (6 to 18 years) at a local science museum. Factor analysis was conducted to identify OC trait dimensions. Univariate and multivariate twin modeling was performed to estimate the heritability of OC trait dimensions in a subset of twins (220 pairs).\n\nResultsSix OC dimensions were identified: Cleaning/Contamination, Hoarding, Rumination, Superstition, Counting/Checking, and Symmetry/Ordering. The TOCS total score (74%) and OC trait dimensions were heritable (30-77%). Hoarding was phenotypically distinct but shared genetic effects with other OC dimensions. Most of the genetic effects were shared between dimensions while unique environment accounted for the majority of dimension-specific variance, except for hoarding which had considerable unique genetic factors. A latent trait did not account for the shared variance between dimensions.\n\nConclusionsOC traits and individual OC dimensions were heritable, although the degree of shared and dimension-specific etiological factors varied by dimension. The TOCS is useful for genetic research of OC traits and OC dimensions should be examined individually and together along with total trait scores to characterize OC genetic architecture.

genetics

Genome-wide association study meta-analysis of the Alcohol Use Disorder Identification Test (AUDIT) in two population-based cohorts (N=141,958)

Alcohol use disorders (AUD) are common conditions that have enormous social and economic consequences. We obtained quantitative measures using the Alcohol Use Disorder Identification Test (AUDIT) from two population-based cohorts of European ancestry: UK Biobank (UKB; N=121,604) and 23andMe (N=20,328) and performed a genome-wide association study (GWAS) meta-analysis. We also performed GWAS for AUDIT items 1-3, which focus on consumption (AUDIT-C), and for items 4-10, which focus on the problematic consequences of drinking (AUDIT-P). The GWAS meta-analysis of AUDIT total score identified 10 associated risk loci. Novel associations localized to genes including JCAD and SLC39A13; we also replicated previously identified signals in the genes ADH1B, ADH1C, KLB, and GCKR. The dimensions of AUDIT showed positive genetic correlations with alcohol consumption (rg=0.76-0.92) and Diagnostic and Statistical Manual of Mental Disorders (DSM-IV) alcohol dependence (rg=0.33-0.63). AUDIT-P and AUDIT-C showed significantly different patterns of association across a number of traits, including psychiatric disorders. AUDIT-P was positively genetically correlated with schizophrenia (rg=0.22, p=3.0x10-10), major depressive disorder (rg=0.26, p=5.6x10-3), and attention-deficit/hyperactivity disorder (ADHD; rg=0.23, p=1.1x10-5), whereas AUDIT-C was negatively genetically correlated with major depressive disorder (rg=-0.24, p=3.7x10-3) and ADHD (rg=-0.10, p=1.8x10-2). We also used the AUDIT data in the UKB to identify thresholds for dichotomizing AUDIT total score that optimize genetic correlations with DSM-IV alcohol dependence. Coding individuals with AUDIT total score of [≤]4 as controls and [≥]12 as cases produced a high genetic correlation with DSM-IV alcohol dependence (rg=0.82, p=3.2x10-6) while retaining most subjects. We conclude that AUDIT scores ascertained in population-based cohorts can be used to explore the genetic basis of both alcohol consumption and AUD.

genetics

The Neurospora crassa standard Oak Ridge background exhibits an atypically efficient meiotic silencing by unpaired DNA.

Meiotic silencing by unpaired DNA (MSUD) was discovered in crosses made in the standard Oak Ridge (OR) genetic background of Neurospora crassa. However, MSUD often was decidedly less efficient when the OR-derived MSUD tester strains were crossed with wild-isolated strains (W), which suggested either that sequence heterozygosity in tester x W crosses suppresses MSUD, or that OR represents the MSUD-conducive extreme in the range of genetic variation in MSUD efficiency. Our results support the latter model. MSUD was much less efficient in near-isogenic crosses made in a novel N. crassa B/S1 and the N. tetrasperma 85 genetic backgrounds. Possibly, regulatory cues that in other genetic backgrounds calibrate the MSUD response are missing from OR. The OR versus B/S1 difference appears to be determined by loci on chromosomes 1, 2, and 5. OR crosses heterozygous for a duplicated chromosome segment (Dp) have for long been known to exhibit an MSUD-dependent barren phenotype. However, inefficient MSUD in N. tetrasperma 85 made Dp-heterozygous crosses non-barren. This is germane to our earlier demonstration that Dps can act as dominant suppressors of repeat-induced point mutation (RIP). Occasionally, during ascospore partitioning rare asci contained >8 nuclei, and round ascospores dispersed less efficiently than spindle-shaped ones.\n\nGeneral abstractIn crosses made in the standard OR genetic background of Neurospora crassa, an RNAi-mediated process called MSUD efficiently silences any gene not properly paired with its homologue during meiosis. We found that MSUD was not as efficient in comparable crosses made in the N. crassa B/S1 and N. tetrasperma 85 backgrounds, suggesting that efficient MSUD is not necessarily the norm in Neurospora. Indeed, using OR strains for genetic studies probably fortuitously facilitated the discovery of MSUD and its suppressors. As few as three unlinked loci appear to underlie the OR versus B/S1 difference in MSUD.

genetics