Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Genetic conflict with a parasitic nematode disrupts the legume-rhizobia mutualism

Genetic variation for partner quality in mutualisms is an evolutionary paradox. One possible resolution to this puzzle is that there is a tradeoff between partner quality and other fitness-related traits. Here, we tested whether a susceptibility to parasitism is one such tradeoff in the mutualism between legumes and nitrogen-fixing bacteria (rhizobia). We performed two greenhouse experiments with the legume Medicago truncatula. In the first, we inoculated each plant with the rhizobia Ensifer meliloti and with one of 40 genotypes of the parasitic root-knot nematode Meloidogyne hapla. In the second experiment, we inoculated all plants with rhizobia and half of the plants with a genetically variable population of nematodes. Using the number of nematode galls as a proxy for infection severity, we found that plant genotypes differed in susceptibility to nematode infection, and nematode genotypes differed in infectivity. Second, we showed that there was a genetic correlation between the number of mutualistic structures formed by rhizobia (nodules) and the number of parasitic structures formed by nematodes (galls). Finally, we found that nematodes disrupt the rhizobia mutualism: nematode-infected plants formed fewer nodules and had less nodule biomass than uninfected plants. Our results demonstrate that there is genetic conflict between attracting rhizobia and repelling nematodes in Medicago. If genetic conflict with parasitism is a general feature of mutualism, it could account for the maintenance of genetic variation in partner quality and influence the evolutionary dynamics of positive species interactions.\n\nImpact summaryCooperative species interactions, known as mutualisms, are vital for organisms from plants to humans. For example, beneficial microbes in the human gut are a necessary component of digestive health. However, parasites often infect their hosts via mechanisms that are extraordinarily similar to those used by mutualists, which may create a tradeoff between attracting mutualists and resisting parasites. In this study, we investigated whether this tradeoff exists, and how parasites impact mutualism function in the barrelclover Medicago truncatula, a close relative of alfalfa. Legumes like Medicago depend on nitrogen provided by mutualistic bacteria (rhizobia) to grow, but they are also infected by parasitic worms called nematodes, which steal plant nutrients. Both microorganisms live in unique structures (nodules and galls) on plant roots. We showed that the benefits of mutualism and the costs of parasitism are predicted by the number of mutualistic structures (nodules) and the number of parasitic structures (galls), respectively. Second, we found that there is a genetic tradeoff between attracting mutualists and repelling parasites in Medicago truncatula: plant genotypes that formed more rhizobia nodules also formed more nematode galls. Finally, we found that nematodes disrupt the rhizobia mutualism. Nematode-infected plants formed fewer rhizobia nodules and less total nodule biomass than uninfected plants. Our research addresses an enduring evolutionary puzzle: why is there so much variation in the benefits provided by mutualists when natural selection should weed out low-quality partners? Tradeoffs between benefits provided by mutualists and their susceptibility to parasites could resolve this paradox.

evolutionary biology

The standard genetic code facilitates exploration of the space of functional nucleotide sequences

The standard genetic code is well known to be optimized for minimizing the phenotypic effects of single nucleotide substitutions, a property that was likely selected for during the emergence of a universal code. Given the fitness advantage afforded by high standing genetic diversity in a population in a dynamic environment, it is possible that selection to explore a large fraction of the space of functional proteins also occurred. To determine whether selection for such a property played a role during the emergence of the nearly universal genetic code, we investigated the number of functional variants of the Escherichia coli PhoQ protein explored at different time scales under translation using different genetic codes. We found that the standard genetic code is highly optimal for exploring a large fraction of the space of functional PhoQ variants at intermediate time scales as compared to random codes. Environmental changes, in response to which genetic diversity in a population provides a fitness advantage, are likely to have occurred at these intermediate time scales. Our results indicate that the ability of the standard code to explore a large fraction of the space of functional sequence variants arises from a balance between robustness and flexibility and is largely independent of the property of the standard code to minimize the phenotypic effects of mutations. We propose that selection to explore a large fraction of the functional sequence space while minimizing the phenotypic effects of mutations contributed towards the emergence of the standard code as the universal genetic code.

evolutionary biology

Genetic architecture of gene expression traits across diverse populations

For many complex traits, gene regulation is likely to play a crucial mechanistic role. How the genetic architectures of complex traits vary between populations and subsequent effects on genetic prediction are not well understood, in part due to the historical paucity of GWAS in populations of non-European ancestry. We used data from the MESA (Multi-Ethnic Study of Atherosclerosis) cohort to characterize the genetic architecture of gene expression within and between diverse populations. Genotype and monocyte gene expression were available in individuals with African American (AFA, n=233), Hispanic (HIS, n=352), and European (CAU, n=578) ancestry. We performed expression quantitative trait loci (eQTL) mapping in each population and show genetic correlation of gene expression depends on shared ancestry proportions. Using elastic net modeling with cross validation to optimize genotypic predictors of gene expression in each population, we show the genetic architecture of gene expression for most predictable genes is sparse. We found the best predicted gene, TACSTD2, was the same across populations with R2 > 0.86 in each population. However, we identified a subset of genes that are well-predicted in one population, but poorly predicted in another. We show these differences in predictive performance are due to allele frequency differences between populations. Using genotype weights trained in MESA to predict gene expression in independent populations showed that a training set with ancestry similar to the test set is better at predicting gene expression in test populations, demonstrating an urgent need for diverse population sampling in genomics. Our predictive models and performance statistics in diverse cohorts are made publicly available for use in transcriptome mapping methods at .\n\nAuthor summaryMost genome-wide association studies (GWAS) have been conducted in populations of European ancestry leading to a disparity in understanding the genetics of complex traits between populations. For many complex traits, gene regulation is critical, given the consistent enrichment of regulatory variants among trait-associated variants. However, it is still unknown how the effects of these key variants differ across populations. We used data from MESA to study the underlying genetic architecture of gene expression by optimizing gene expression prediction within and across diverse populations. The populations with genotype and gene expression data available are from individuals with African American (AFA, n=233), Hispanic (HIS, n=352), and European (CAU, n=578) ancestry. After calculating the prediction performance, we found that there are many genes that were well predicted in one population are poorly predicted in another. We further show that a training set with ancestry similar to the test set resulted in better gene expression predictions, demonstrating the need to incorporate diverse populations in genomic studies. Our gene expression prediction models and performance statistics are publicly available to facilitate future transcriptome mapping studies in diverse populations.

genomics

A meta-analysis of the diagnostic sensitivity and clinical utility of genome sequencing, exome sequencing and chromosomal microarray in children with suspected genetic diseases

IMPORTANCEGenetic diseases are a leading cause of childhood mortality. Whole genome sequencing (WGS) and whole exome sequencing (WES) are relatively new methods for diagnosing genetic diseases.\n\nOBJECTIVESCompare the diagnostic sensitivity (rate of causative, pathogenic or likely pathogenic genotypes in known disease genes) and rate of clinical utility (proportion in whom medical or surgical management was changed by diagnosis) of WGS, WES, and chromosomal microarrays (CMA) in children with suspected genetic diseases.\n\nDATA SOURCES AND STUDY SELECTIONSystematic review of the literature (January 2011 - August 2017) for studies of diagnostic sensitivity and/or clinical utility of WGS, WES, and/or CMA in children with suspected genetic diseases. 2% of identified studies met selection criteria.\n\nDATA EXTRACTION AND SYNTHESISTwo investigators extracted data independently following MOOSE/PRISMA guidelines.\n\nMAIN OUTCOMES AND MEASURESPooled rates and 95% Cl were estimated with a random-effects model. Metaanalysis of the rate of diagnosis was based on test type, family structure, and site of testing.\n\nRESULTSIn 36 observational series and one randomized control trial, comprising 20,068 children, the diagnostic sensitivity of WGS (0.41, 95% Cl 0.34-0.48, I2=44%) and WES (0.35, 95% Cl 0.31-0.39, I2=85%) were qualitatively greater than CMA (0.10, 95% Cl 0.08-0.12, I2=81%). Subgroup meta-analyses showed that the diagnostic sensitivity of WGS was significantly greater than CMA in studies published in 2017 (P<.0001, I2=13% and I2=40%, respectively), and the diagnostic sensitivity of WES was significantly greater than CMA in studies featuring within-cohort comparisons (P<001, I2=36%). Evidence for a significant difference in the diagnostic sensitivity of WGS and WES was lacking. In studies featuring within-cohort comparisons of singleton and trio WGS/WES, the likelihood of diagnosis was significantly greater for trios (odds ratio 2.04, 95% Cl 1.62-2.56, I2=12%; P<.0001). The diagnostic sensitivity of WGS/WES with hospital-based interpretation (0.41, 95% Cl 0.38-0.45, I2=50%) was qualitatively higher than that of reference laboratories (0.28, 95% Cl 0.24-0.32, I2=81%); this difference was significant in meta-analysis of studies published in 2017 (P=.004, I2=34% and I2=26%, respectively). The rates of clinical utility of WGS (0.27, 95% Cl 0.17-0.40, I2=54%) and WES (0.18, 95% Cl 0.13-0.24, I2-77%) were higher than CMA (0.06, 95% Cl 0.05-0.07, I2=42%); this difference was significant in meta-analysis of WGS vs CMA (P<.0001).\n\nCONCLUSIONS AND RELEVANCEIn children with suspected genetic diseases, the diagnostic sensitivity and rate of clinical utility of WGS/WES were greater than CMA. Subgroups with higher WGS/WES diagnostic sensitivity were trios and those receiving hospital-based interpretation. WGS/WES should be considered a first-line genomic test for children with suspected genetic diseases.\n\nKey PointsO_ST_ABSQuestionC_ST_ABSWhat is the relative diagnostic sensitivity and clinical utility of different genome tests in children with suspected genetic diseases?\n\nFindingsWhole genome sequencing had greater diagnostic sensitivity and clinical utility than chromosomal microarrays. Testing parent-child trios had greater diagnostic sensitivity than proband singletons. Hospital-based testing had greater diagnostic sensitivity than reference laboratories.\n\nMeaningTrio genomic sequencing is the most sensitive diagnostic test for children with suspected genetic diseases.

genomics

Genetic basis of thermal plasticity variation in Drosophila melanogaster body size

Body size is a quantitative trait that is closely associated to fitness and under the control of both genetic and environmental factors. While developmental plasticity for this and other traits is heritable and under selection, little is known about the genetic basis for variation in plasticity that can provide the raw material for its evolution. We quantified genetic variation for body size plasticity in Drosophila melanogaster by measuring thorax and abdomen length of females reared at two temperatures from a panel representing naturally segregating alleles, the Drosophila Genetic Reference Panel (DGRP). We found variation between genotypes for the levels and direction of thermal plasticity in size of both body parts. We then used a Genome-Wide Association Study (GWAS) approach to unravel the genetic basis of inter-genotype variation in body size plasticity, and used different approaches to validate selected QTLs and to explore potential pleiotropic effects. We found mostly \"private QTLs\", with little overlap between the candidate loci underlying variation in plasticity for thorax versus abdomen size, for different properties of the plastic response, and for size versus size plasticity. We also found that the putative functions of plasticity QTLs were diverse and that alleles for higher plasticity were found at lower frequencies in the target population. Importantly, a number of our plasticity QTLs have been targets of selection in other populations. Our data sheds light onto the genetic basis of inter-genotype variation in size plasticity that is necessary for its evolution.\n\nSignificance StatementThe environmental conditions under which development takes place can affect developmental outcomes and lead to the production of phenotypes adjusted to the environment adults will live in. This developmental plasticity, which can help organisms cope with environmental heterogeneity, is heritable and under selection. Plasticity can itself evolve, a process that will be partly dependent on the available genetic variation for this trait. Using a wild-derived D. melanogaster panel, we identified DNA sequence variants associated to variation in thermal plasticity for body size. We found that these variants correspond to a diverse set of gene functions. Furthermore, their effects differ between body parts and properties of the thermal response, which can, therefore, evolve independently. Our results shed new light onto a number of key questions about the long discussed genes for plasticity.

evolutionary biology

Engineered Genetic Redundancy Relaxes Selective Constraints upon Endogenous Genes in Viral RNA Genomes

1Genetic redundancy, understood as the functional overlap of different genes, is a double-edge sword. At the one side, it is thought to serve as a robustness mechanism that buffers the deleterious effect of mutations hitting one of the redundant copies, thus resulting in pseudogenization. At the other side, it is considered as a source of genetic and functional innovation. In any case, genetically redundant genes are expected to show an acceleration in the rate of molecular evolution. Here we tackle the role of genetic redundancy in viral RNA genomes. To this end, we have evaluated the rates of compensatory evolution for deleterious mutations affecting an essential function, the suppression of RNA silencing plant defense, of tobacco etch potyvirus (TEV). TEV genotypes containing deleterious mutations in presence/absence of engineered genetic redundancy were evolved and the pattern of fitness and virulence recovery evaluated. Genetically redundant genotypes suffered less from the effect of deleterious mutations and showed relatively minor changes in fitness and virulence. By contrast, non-genetically redundant genotypes had very low fitness and virulence at the beginning of the evolution experiment that were fully recovered by the end. At the molecular level, the outcome depended on the combination of the actual mutations being compensated and the presence/absence of genetic redundancy. Reversions to wild-type alleles were the norm in the non-redundant genotypes while redundant ones either did not fix any mutation at all or showed a higher nonsynonymous mutational load.

evolutionary biology

The structure of the genetic code as an optimal graph clustering problem

The standard genetic code (SGC) is the set of rules by which genetic information is translated into proteins, from codons, i.e. triplets of nucleotides, to amino acids. The questions about the origin and the main factor responsible for the present structure of the code are still under a hot debate. Various methodologies have been used to study the features of the code and assess the level of its potential optimality. Here, we introduced a new general approach to evaluate the quality of the genetic code structure. This methodology comes from graph theory and allows us to describe new properties of the genetic code in terms of conductance. This parameter measures the robustness of codon groups against the potential changes in translation of the protein-coding sequences generated by single nucleotide substitutions. We described the genetic code as a partition of an undirected and unweighted graph, which makes the model general and universal. Using this approach, we showed that the structure of the genetic code is a solution to the graph clustering problem. We presented and discussed the structure of the codes that are optimal according to the conductance. Despite the fact that the standard genetic code is far from being optimal according to the conductance, its structure is characterised by many codon groups reaching the minimum conductance for their size. The SGC represents most likely a local minimum in terms of errors occurring in protein-coding sequences and their translation.

systems biology

Concordance Of Genetic Variation That Increases Risk For Tourette Syndrome And That Influences Its Underlying Neurocircuitry

BACKGROUNDThere have been considerable recent advances in understanding the genetic architecture of Tourette Syndrome (TS) as well as its underlying neurocircuitry. However, the mechanisms by which genetic variations that increase risk for TS - and its main symptom dimensions - influence relevant brain regions are poorly understood. Here we undertook a genome-wide investigation of the overlap between TS genetic risk and genetic influences on the volume of specific subcortical brain structures that have been implicated in TS.\n\nMETHODSWe obtained summary statistics for the most recent TS genome-wide association study (GWAS) from the TS Psychiatric Genomics Consortium Working Group (4,644 cases and 8,695 controls) and GWAS of subcortical volumes from the ENIGMA consortium (30,717 individuals). We also undertook analyses using GWAS summary statistics of key symptom factors in TS, namely social disinhibition and symmetry behaviour. SNP Effect Concordance Analysis (SECA) was used to examine genetic pleiotropy - the same SNP affecting two traits - and concordance - the agreement in SNP effect directions across these two traits. In addition, a conditional false discovery rate (FDR) analysis was performed, conditioning the TS risk variants on each of the seven subcortical and the intracranial brain volume GWAS. Linkage Disequilibrium Score Regression (LDSR) was used as validation of SECA.\n\nRESULTSSECA revealed significant pleiotropy between TS and putaminal (p=2x10-4) and caudal (p=4x10-4) volumes, independent of direction of effect, and significant concordance between TS and lower thalamic volume (p=1x10-3). LDSR lent additional support for the association between TS and thalamic volume (p=5.85x10-2). Furthermore, SECA revealed significant evidence of concordance between the social disinhibition symptom dimension and lower thalamic volume (p=1x10-3), as well as concordance between symmetry behaviour and greater putaminal volume (p=7x10-4). Conditional FDR analysis further revealed novel variants significantly associated with TS (p<8x10-7) when conditioning on intracranial (rs2708146, q=0.046; and rs72853320, q=0.035 and hippocampal (rs1922786, q=0.001 volumes respectively.\n\nCONCLUSIONThese data indicate concordance for genetic variations involved in disorder risk and subcortical brain volumes in TS. Further work with larger samples is needed to fully delineate the genetic architecture of these disorders and their underlying neurocircuitry.

genomics

Genetic characterization of chytrids isolated from larval amphibians collected in central and east Texas

Chytridiomycosis, an emerging infectious disease caused by the fungal pathogen Batrachochytrium dendrobatidis (Bd), has caused amphibian population declines worldwide. Bd was first described in the 1990s and there are still geographic gaps in the genetic analysis of this globally distributed pathogen. Relatively few genetic studies have focused on regions where Bd exhibits low virulence, potentially creating a bias in our current knowledge of the pathogens genetic diversity. Disease-associated declines have not been recorded in Texas (USA), yet Bd has been detected on amphibians in the state. These strains have not been isolated and characterized genetically; therefore, we isolated, cultured, and genotyped Bd from central Texas and compared isolates to a panel of previously genotyped strains distributed across the Western Hemisphere. We also isolated other chytrids not known to infect amphibians from east Texas. To identify larval amphibian hosts, we sequenced part of the COI gene. Among 37 Bd isolates from Texas, we detected 19 unique multi-locus genotypes, but found no genetic structure associated with host species, Texas localities, or across North America. Isolates from central Texas exhibit high diversity and genetically cluster with Bd-GPL isolates from the western U.S. that have caused amphibian population declines. This study genetically characterizes isolates of Bd from the south central U.S. and adds to the global knowledge of Bd genotypes.

evolutionary biology

Comparative genetic architectures of schizophrenia in East Asian and European populations

Author summarySchizophrenia is a severe psychiatric disorder with a lifetime risk of about 1% world-wide. Most large schizophrenia genetic studies have studied people of primarily European ancestry, potentially missing important biological insights. Here we present a study of East Asian participants (22,778 schizophrenia cases and 35,362 controls), identifying 21 genome-wide significant schizophrenia associations in 19 genetic loci. Over the genome, the common genetic variants that confer risk for schizophrenia have highly similar effects in those of East Asian and European ancestry (rg=0.98), indicating for the first time that the genetic basis of schizophrenia and its biology are broadly shared across these world populations. A fixed-effect meta-analysis including individuals from East Asian and European ancestries revealed 208 genome-wide significant schizophrenia associations in 176 genetic loci (53 novel). Trans-ancestry fine-mapping more precisely isolated schizophrenia causal alleles in 70% of these loci. Despite consistent genetic effects across populations, polygenic risk models trained in one population have reduced performance in the other, highlighting the importance of including all major ancestral groups with sufficient sample size to ensure the findings have maximum relevance for all populations.

genetics

Genetic dissection of MAPK-mediated complex traits across S. cerevisiae

Signaling pathways enable cells to sense and respond to their environment. Many cellular signaling strategies are conserved from fungi to humans, yet their activity and phenotypic consequences can vary extensively among individuals within a species. A systematic assessment of the impact of naturally occurring genetic variation on signaling pathways remains to be conducted. In S. cerevisiae, both response and resistance to stressors that activate signaling pathways differ between diverse isolates. Here, we present a quantitative trait locus (QTL) mapping approach that enables us to identify genetic variants underlying such phenotypic differences across the genetic and phenotypic diversity of S. cerevisiae. Using a Round-robin cross between twelve diverse strains, we determined the genetic architectures of phenotypes critically dependent on MAPK signaling cascades. Genetic variants identified fell within MAPK signaling networks themselves as well as other interconnected signaling pathways, illustrating how genetic variation can shape the phenotypic output of highly conserved signaling cascades.

Genetics

Inference and analysis of population structure using genetic data and network theory

Clustering individuals to subpopulations based on genetic data has become commonplace in many genetic studies. Inference of population structure is most often done by applying model-based approaches, aided by visualization using distance-based approaches such as multidimensional scaling. While existing distance-based approaches suffer from lack of statistical rigor, model-based approaches entail assumptions of prior conditions such as that the subpopulations are at Hardy-Weinberg equilibria. Here we present a distance-based approach for inference of population structure using genetic data by defining population structure using network theory terminology and methods. A network is constructed from a pairwise genetic-similarity matrix of all sampled individuals. The community partition, a partition of a network to dense subgraphs, is equated with population structure, a partition of the population to genetically related groups. Community detection algorithms are used to partition the network into communities, interpreted as a partition of the population to subpopulations. The statistical significance of the structure can be estimated by using permutation tests to evaluate the significance of the partitions modularity, a network theory measure indicating the quality of community partitions. In order to further characterize population structure, a new measure of the Strength of Association (SA) for an individual to its assigned community is presented. The Strength of Association Distribution (SAD) of the communities is analyzed to provide additional population structure characteristics, such as the relative amount of gene flow experienced by the different subpopulations and identification of hybrid individuals. Human genetic data and simulations are used to demonstrate the applicability of the analyses. The approach presented here provides a novel, computationally efficient, model-free method for inference of population structure which does not entail assumption of prior conditions. The method is implemented in the software NetStruct, available at https://github.com/GiliG/NetStruct.

Genetics

Genetic evidence that lower circulating FSH levels lengthen menstrual cycle, increase age at menopause, and impact reproductive health: a UK Biobank study

Study question: How does a genetic variant altering follicle stimulating hormone (FSH) levels, which we identified as associated with length of menstrual cycle, more widely impact reproductive health?\n\nSummary answer: The T allele of the FSHB promoter polymorphism (rs10835638) results in longer menstrual cycles and later menopause and, while having detrimental effects on fertility, is protective against endometriosis.\n\nWhat is known already: The FSHB promoter polymorphism (rs10835638) affects levels of FSHB transcription and, as a result, levels of FSH. FSH is required for normal fertility and genetic variants at the FSHB locus are associated with age at menopause and polycystic ovary syndrome (PCOS).\n\nStudy design, size, duration: We conducted a genetic association study using cross-sectional data from the UK Biobank.\n\nParticipants/materials, setting, methods: We included white British individuals aged 40-69 years in 2006-2010, included in the May 2015 release of genetic data from UK Biobank. We conducted a genome-wide association study (GWAS) in 9,534 individuals to identify genetic variants associated with length of menstrual cycle. We tested the FSH lowering T allele of the FSHB promoter polymorphism (rs10835638) for associations with 29 reproductive phenotypes in up to 63,350 individuals.\n\nMain results and the role of chance: In the GWAS for menstrual cycle length, only variants near the FSHB gene reached genome-wide significance (P<5x10-8). The FSH-lowering T allele of the FSHB promoter polymorphism (rs10835638G>T; MAF 0.16) was associated with longer menstrual cycles (0.16 s.d. (approx. 1 day) per minor allele; 95% CI 0.12-0.20; P=6x10-16), later age at menopause (0.13 years per minor allele; 95% CI 0.04-0.22; P=5.7x10-3), greater female nulliparity (OR=1.06; 95% CI 1.02-1.11; P=4.8x10-3) and lower risk of endometriosis (OR=0.79; 95% CI 0.69-0.90; P=4.1x10-4). The FSH-lowering T allele was not associated more generally with other reproductive illnesses or conditions and we did not replicate associations with male infertility or PCOS.\n\nLimitations, reasons for caution: The data included might be affected by recall bias. Women with a cycle length recorded were aged over 40 and were approaching menopause, however we did not find evidence that this affected the results. Many of the illnesses had relatively small sample sizes and so we may have been under-powered to detect an effect.\n\nWider implications of the findings: We found a strong novel association between a genetic variant that lowers FSH levels and longer menstrual cycles, at a locus previously robustly associated with age at menopause. The variant was also associated with nulliparity and endometriosis risk. We conclude that lifetime differences in circulating levels of FSH between individuals can influence menstrual cycle length and a range of reproductive outcomes, including menopause timing, infertility, endometriosis and PCOS.

Genetics

Genetic heterogeneity in autism: from single gene to a pathway perspective

AbstractThe extreme genetic heterogeneity of autism spectrum disorder (ASD) represents a major challenge. Recent advances in genetic screening and systems biology approaches have extended our knowledge of the genetic etiology of ASD. In this review, we discuss the paradigm shift from a single gene causation model to pathway perturbation model as a guide to better understand the pathophysiology of ASD. We discuss recent genetic findings obtained through next-generation sequencing (NGS) and examine various integrative analyses using systems biology and complex networks approaches that identify convergent patterns of genetic elements associated with ASD. This review provides a summary of the genetic findings of family-based genome screening studies.

Genetics

Mortality Selection in a Genetic Sample and Implications for Association Studies

Mortality selection is a general concern in the social and health sciences. Recently, existing health and social science cohorts have begun to collect genomic data. Causes of selection into a genomic dataset can influence results from genomic analyses. Selective non-participation, which is specific to a particular study and its participants, has received attention in the literature. But mortality selection--the very general phenomenon that genomic data collected at a particular age represents selective participation by only the subset of birth cohort members who have survived to the time of data collection--has been largely ignored. Here we test the hypothesis that such mortality selection may significantly alter estimates in polygenetic association studies of both health and non-health traits. We demonstrate mortality selection into genome-wide SNP data collection at older ages using the U.S.-based Health and Retirement Study (HRS). We then model the selection process. Finally, we test whether mortality selection alters estimates from genetic association studies. We find evidence for mortality selection. Healthier and more socioeconomically advantaged individuals are more likely to survive to be eligible to participate in the genetic sample of the HRS. Mortality selection leads to modest drift in estimating time-varying genetic effects, a drift that is enhanced when estimates are produced from data that has additional mortality selection. There is no general solution for correcting for mortality selection in a birth cohort prior to entry into a longitudinal study. We illustrate how genetic association studies using HRS data can adjust for mortality selection from study entry to time of genetic data collection by including probability weights that account for mortality selection. Mortality selection should be investigated more broadly in genetically-informed samples from other cohort studies.

Genetics

Associations of coffee genetic risk scores with coffee, tea and other beverages in the UK Biobank

BackgroundGenetic variants which determine amount of coffee consumed have been identified in genome-wide association studies (GWAS) of coffee consumption; these may help to further understanding of the effects of coffee on health outcomes. However, there is limited information about how these variants relate to caffeinated beverage consumption more generally.\n\nAimsTo improve phenotype definition for coffee consumption related genetic risk scores by testing their association with coffee, tea and other beverages.\n\nMethodsWe tested the associations of genetic risk scores for coffee consumption with beverage consumption in 114,316 individuals of European ancestry from the UK Biobank. Drinks were self-reported in a baseline questionnaire and in detailed 24 dietary recall questionnaires in a subset.\n\nResultsGenetic risk scores including two and eight single nucleotide polymorphisms (SNPs) explained up to 0.39%, 0.19% and 0.77% of the variance in coffee, tea and combined coffee and tea consumption respectively. A one standard deviation increase in the 8 SNP genetic risk score was associated with a 0.13 cup per day (95% CI: 0.12, 0.14), 0.12 cup per day (95%CI: 0.11, 0.14) and 0.25 cup per day (95% CI: 0.24, 0.27) increase in coffee, tea and combined tea and coffee consumption, respectively. Genetic risk scores also demonstrated positive associations with both caffeinated and decaffeinated coffee and tea consumption. In 48,692 individuals with dietary recall data, the genetic risk scores were positively associated with coffee and tea, (apart from herbal teas) consumption, but did not show clear evidence for positive associations with other beverages. However, there was evidence that the genetic risk scores were associated with lower daily water consumption and lower overall drink consumption.\n\nConclusionsGenetic risk scores created from variants identified in coffee consumption GWAS associate more broadly with caffeinated beverage consumption and also with decaffeinated coffee and tea consumption.

genetics

Joint genetic analysis using variant sets reveals polygenic gene-context interactions

Joint genetic models for multiple traits have helped to enhance association analyses. Most existing multi-trait models have been designed to increase power for detecting associations, whereas the analysis of interactions has received considerably less attention. Here, we propose iSet, a method based on linear mixed models to test for interactions between sets of variants and environmental states or other contexts. Our model generalizes previous interaction tests and in particular provides a test for local differences in the genetic architecture between contexts. We first use simulations to validate iSet before applying the model to the analysis of genotype-environment interactions in an eQTL study. Our model retrieves a larger number of interactions than alternative methods and reveals that up to 20% of cases show context-specific configurations of causal variants. Finally, we apply iSet to test for sub-group specific genetic effects in human lipid levels in a large human cohort, where we identify a gene-sex interaction for C-reactive protein that is missed by alternative methods.\n\nAuthor summaryGenetic effects on phenotypes can depend on external contexts, including environment. Statistical tests for identifying such interactions are important to understand how individual genetic variants may act in different contexts. Interaction effects can either be studied using measurements of a given phenotype in different contexts, under the same genetic backgrounds, or by stratifying a population into subgroups. Here, we derive a method based on linear mixed models that can be applied to both of these designs. iSet enables testing for interactions between context and sets of variants, and accounts for polygenic effects. We validate our model using simulations, before applying it to the genetic analysis of gene expression studies and genome-wide association studies of human blood lipid levels. We find that modeling interactions with variant sets offers increased power, thereby uncovering interactions that cannot be detected by alternative methods.

genetics

Bayesian analysis of genetic association across tree-structured routine healthcare data in the UK Biobank

Genetic discovery from the multitude of phenotypes extractable from routine healthcare data has the ability to radically transform our understanding of the human phenome, thereby accelerating progress towards precision medicine. However, a critical question when analysing high-dimensional and heterogeneous data is how to interrogate increasingly specific subphenotypes whilst retaining statistical power to detect genetic associations. Here we develop and employ a novel Bayesian analysis framework that exploits the hierarchical structure of diagnosis classifications to jointly analyse genetic variants against UK Biobank healthcare phenotypes. Our method displays a more than 20% increase in power to detect genetic effects over other approaches, such that we uncover the broader burden of genetic variation: we identify associations with over 2,000 diagnostic terms. We find novel associations with common immune-mediated diseases (IMD), we reveal the extent of genetic sharing between specific IMDs, and we expose differences in disease perception or diagnosis with potential clinical implications.

genetics