Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

Education and coronary heart disease: a Mendelian randomization study

ObjectivesTo determine whether educational attainment is a causal risk factor in the development of coronary heart disease.\n\nDesignMendelian randomization study, where genetic data are used as proxies for education, in order to minimize confounding. A two-sample design was applied, where summary level genetic data was analysed from two publically available consortia.\n\nSettingIn the main analysis, we analysed genetic data from two large consortia (CARDIoGRAM and SSGAC), comprising of 112 cohorts from predominantly high-income countries. In addition, we also analysed genetic data from 7 additional large consortia, in order to identify putative causal mediators.\n\nParticipantsThe main analysis was of 589 377 men and women, predominantly of European origin.\n\nExposureA one standard deviation increase in the genetic predisposition towards higher education (i.e. 3.6 years of additional schooling). This was measured by 162 genetic variants that have been previously associated with education.\n\nMain outcomeCombined fatal and nonfatal coronary heart disease (63 746 events).\n\nResults3.6 years of additional education lowered the risk of coronary heart disease by a third (odds ratio = 0.67, 95% confidence interval [CI], 0.59 to 0.77, p=0.01). Equivalent increases in education were also causally associated with reductions in smoking, BMI and improvements in blood lipid profiles.\n\nConclusionsMore time spent in education is causally associated with a large reduction in the risk of coronary heart disease. This may be partly explained by changes to smoking, BMI and a blood lipids. These findings offer support for policy interventions that increase education, in order to also reduce the burden of cardiovascular disease.

epidemiology

On the heritability of criminal justice processing

An impressive number of researchers have devoted a great amount of effort toward examining various predictors of criminal justice processing outcomes. Indeed, a vast amount of research has examined various individual- and aggregate-level predictors of arrests, incarceration, and sentencing decisions. To this point, less attention has been devoted toward uncovering the relative contribution of genetic and environmental effects on variation in risk for criminal justice processing. As a result, the current study employs a behavioral genetic design in order to help fill this void in the existing literature. Using twin data from a national sample of youth, the current study produced evidence suggesting that genetic factors accounted for at least a portion of variance in risk for incarceration among female twins and probation among male twins. Shared and nonshared environmental influences accounted for the variance in risk for arrest among both female and male twins, probation among female twins, and incarceration among male twins. Ultimately, it appears that risk for contact with the criminal justice system and criminal justice processing is structured by a combination of factors beyond shared cultural and neighborhood environments, and appear to also include genetic factors as well. Moving forward, continuing to not use genetically sensitive research designs capable of estimating the role of genetic and nonshared environmental influences on criminal justice outcomes may result in misleading results.

epidemiology

Efficient And Accurate Causal Inference With Hidden Confounders From Genome-Transcriptome Variation Data

Mapping gene expression as a quantitative trait using whole genome-sequencing and transcriptome analysis allows to discover the functional consequences of genetic variation. We developed a novel method and ultra-fast software Findr for higly accurate causal inference between gene expression traits using cis-regulatory DNA variations as causal anchors, which improves current methods by taking into account hidden confounders and weak regulations. Findr outperformed existing methods on the DREAM5 Systems Genetics challenge and on the prediction of microRNA and transcription factor targets in human lymphoblastoid cells, while being nearly a million times faster. Findr is publicly available at https://github.com/lingfeiwang/findr.\n\nAuthor summaryUnderstanding how genetic variation between individuals determines variation in observable traits or disease risk is one of the core aims of genetics. It is known that genetic variation often affects gene regulatory DNA elements and directly causes variation in expression of nearby genes. This effect in turn cascades down to other genes via the complex pathways and gene interaction networks that ultimately govern how cells operate in an ever changing environment. In theory, when genetic variation and gene expression levels are measured simultaneously in a large number of individuals, the causal effects of genes on each other can be inferred using statistical models similar to those used in randomized controlled trials. We developed a novel method and ultra-fast software Findr which, unlike existing methods, takes into account the complex but unknown network context when predicting causality between specific gene pairs. Findrs predictions have a significantly higher overlap with known gene networks compared to existing methods, using both simulated and real data. Findr is also nearly a million times faster, and hence the only software in its class that can handle modern datasets where the expression levels of ten-thousands of genes are simultaneously measured in hundreds to thousands of individuals.

systems biology

Proof Of Concept: Molecular Prediction Of Schizophrenia Risk

Key PointsO_ST_ABSQuestionC_ST_ABSTo what extent do global polygenic risk scores (PRS), molecular pathway-specific PRS, complement component (C4) gene expression, MHC loci, sex, and ancestry jointly contribute to risk for schizophrenia-spectrum disorders (SZ)?\n\nFindingsGlobal polygenic risk for schizophrenia, sex, and their interaction most robustly predict risk in a classification and regression tree model, with highest risk groups having 50/50 chance of SZ.\n\nMeaningPsychometric risk indicators, such as prodromal symptom assessments, may be enhanced by the examination of genetic risk metrics. Preliminary results suggest that of genetic risk metrics, global polygenic information has the most potential to significantly aide in the prediction of SZ.\n\nAbstractO_ST_ABSImportanceC_ST_ABSSchizophrenia (SZ) has a complex, heterogeneous symptom presentation with limited established associations between biological markers and illness onset. Many (gene) molecular pathways (MPs) are enriched for SZ signal, but it is still unclear how these MPs, global PRS, major histocompatibility complex (MHC) complement component (C4) gene expression, and MHC loci might jointly contribute to SZ and its clinical presentation. It is also unclear whether sex or ancestry interacts with these metrics to increase risk in certain individuals.\n\nObjectiveTo examine multiple genetic metrics, sex, and their interactions as possible predictors of SZ risk. Genetic information could aid in the clinical prediction of risk, but it is still unclear which genetic metrics are most promising, and how sex interacts with genetic risk metrics.\n\nDesign, Setting, and ParticipantsTo examine molecular risk in a proof-of-concept study, we used the Wellcome Trust case-control cohort and classified cases as a function of 1) polygenic risk score (PRS) for both whole genome and for 345 implicated molecular pathways, 2) predicted C4 expression, 3) SZ-relevant MHC loci, 4) sex, and 5) ancestry.\n\nMain Outcomes and MeasuresPRSs, C4 expression, SZ-relevant MHC loci, sex, and ancestry as joint risk factors for SZ.\n\nResultsRecursive partitioning yielded 15 molecular risk classes and retained as significant psychosis classifiers only sex, genome-wide SZ polygenic risk, and one MP PRS. Sex was the most robust classifier in a stepwise regression, and there was a significant interaction of sex with SZ PRS on case status, suggesting males have a lower polygenic risk threshold. By down-sampling case proportion to 1% and 1.4% population base rates in males and females, respectively, high-risk subtypes defined by this model had roughly a 52% odds of developing SZ (individuals with SZ PRS elevated by 2.6 SDs; incidence = 51.8%).\n\nConclusions and RelevanceThis proof-of-concept suggests that global SZ PRS, sex, and their interaction are robust predictors of risk and that males have a lower PRS threshold for onset. Implications for the integration of these metrics with psychometrically-identified risk are discussed.

genomics

The Prevalence And Benefits Of Admixture During Species Invasions: A Role For Epistasis?

Species introductions often bring together genetically divergent source populations, resulting in genetic admixture. This geographic reshuffling of diversity has the potential to generate favorable new genetic combinations, facilitating the establishment and invasive spread of introduced populations. Observational support for the superior performance of admixed introductions has been mixed, however, and the broad importance of admixture to invasion questioned. Under most underlying mechanisms, admixtures benefits should be expected to increasewith greater divergence among and lower genetic diversity within source populations. We use a literature survey to quantify the prevalence of admixture and evaluate whether it occurrs under circumstances predicted to be mostbeneficial to introduced species. We find that 39% of species are reported to be admixed when introduced. Admixed introductions come from sources with a wide range of genetic variation, but are disproportionately absent where there is high genetic divergence among native populations. We discuss multiple potential explanations for these patterns, but note that negative epistatic interactions should be expected at high divergence amongpopulations (outbreeding depression). As a case study, we experimentally cross source populations differing in divergence in the invasive plant Centaurea solstitialis. We find many positive (heterotic) interactions, but fitness benefits decline and are ultimately negative at high source divergence, with patterns suggestingcyto-nuclear epistasis. We conclude that admixture is common in species introductions and often happens under conditions expected to be beneficial to invaders, but that these conditions may be constrained by predictable negativegenetic interactions, potentially explaining conflicting evidence for admixture's benefits to invasion.

evolutionary biology

Human Migration And The Spread Of Malaria Parasites To The New World

BackgroundThe Americas were the last continent to be settled by modern humans, but how and when human malaria parasites arrived in the New World is uncertain. Here, we apply phylogenetic analysis and coalescent-based gene flow modeling to a global collection of Plasmodium falciparum and P. vivax mitogenomes to infer the demographic history and geographic origins of malaria parasites circulating in the Americas. Importantly, we examine P. vivax mitogenomes from previously unsampled forest-covered sites along the Atlantic Coast of Brazil, including the vivax-like species P. simium that locally infects platyrrhini monkeys.\n\nResultsThe best-supported gene flow models are consistent with migration of both malaria parasites from Africa and South Asia to the New World, with no genetic signature of a population bottleneck upon parasite's arrival in the Americas. We found evidence of additional gene flow from Melanesia in P. vivax (but not P. falciparum) mitogenomes from the Americas and speculate that some P. vivax lineages might have arrived with the Australasian peoples who contributed genes to Native Americans in pre-Columbian times. Mitochondrial haplotypes characterized in P. simium from monkeys from the Atlantic Forest are shared by local humans. These vivax-like lineages have not spread to the Amazon Basin, are much less diverse than P. vivax circulating elsewhere in Brazil, and show no close genetic relatedness with P. vivax populations from other continents.\n\nConclusionsEnslaved peoples brought from a wide variety of African locations were major carriers of P. falciparum mitochondrial lineages into the Americas, but additional human migration waves are likely to have contributed to the extensive genetic diversity of present-day New World populations of P. vivax. The reduced genetic diversity of vivax-like monkey parasites, compared with human P. vivax from across this country, argues for a recent human-to-monkey transfer of these lineages in the Atlantic Forest of Brazil.\n\nAuthor summaryMalaria is currently endemic to the Americas, with over 400,000 laboratory-confirmed infections reported annually, but how and when human malaria parasites entered this continent remains largely unknown. To determine the geographic origins of malaria parasites currently circulating in the Americas, we examined a global collection of Plasmodium falciparum and P. vivax mitochondrial genomes, including those from understudied isolates of P. vivax and P. simium, a vivax-like species that infect platyrrhini monkeys, from the Atlantic Forest of Brazil. We found evidence of significant historical migration to the New World of malaria parasites from Africa and, to a lesser extent, South Asia, with further genetic contribution of Melanesian lineages to South American P. vivax populations. Importantly, mitochondrial haplotypes of P. simium are shared by monkeys and humans from the Atlantic Forest, most likely as a result of a recent human-to-monkey transfer. Interestingly, these potentially zoonotic lineages are not found in the Amazon Basin, the main malaria-endemic area in the Americas. We conclude that enslaved Africans were the main carriers of P. falciparum mitochondrial lineages into the Americas, whereas additional migration waves of Australasian peoples and parasites may have contributed to the genetic makeup of present-day New World populations of P. vivax.

microbiology

Fine scale mapping of genomic introgressions within the Drosophila yakuba clade

The process of speciation involves populations diverging over time until they are genetically and reproductively isolated. Hybridization between nascent species was long thought to directly oppose speciation. However, the amount of interspecific genetic exchange (introgression) mediated by hybridization remains largely unknown, although recent progress in genome sequencing has made measuring introgression more tractable. A natural place to look for individuals with admixed ancestry (indicative of introgression) is in regions where species co-occur. In west Africa, D. santomea and D. yakuba hybridize on the island of Sao Tome, while D. yakuba and D. teissieri hybridize on the nearby island of Bioko. In this report, we quantify the genomic extent of introgression between the three species of the Drosophila yakuba clade (D. yakuba, D. santomea), D. teissieri). We sequenced the genomes of 86 individuals from all three species. We also developed and applied a new statistical framework, using a hidden Markov approach, to identify introgression. We found that introgression has occurred between both species pairs but most introgressed segments are small (on the order of a few kilobases). After ruling out the retention of ancestral polymorphism as an explanation for these similar regions, we find that the sizes of introgressed haplotypes indicate that genetic exchange is not recent (>1,000 generations ago). We additionally show that in both cases, introgression was rarer on X chromosomes than on autosomes which is consistent with sex chromosomes playing a large role in reproductive isolation. Even though the two species pairs have stable contemporary hybrid zones, providing the opportunity for ongoing gene flow, our results indicate that genetic exchange between these species is currently rare.\n\nAUTHOR SUMMARYEven though hybridization is thought to be pervasive among animal species, the frequency of introgression, the transfer of genetic material between species, remains largely unknown. In this report we quantify the magnitude and genomic distribution of introgression among three species of Drosophila that encompass the two known stable hybrid zones in this genetic model genus. We obtained whole genome sequences for individuals of the three species across their geographic range (including their hybrid zones) and developed a hidden Markov model-based method to identify patterns of genomic introgression between species. We found that nuclear introgression is rare between both species pairs, suggesting hybrids in nature rarely successfully backcross with parental species. Nevertheless, some D. santomea alleles introgressed into D. yakuba have spread from Sao Tome to other islands in the Gulf of Guinea where D. santomea is not found. Our results indicate that in spite of contemporary hybridization between species that produces fertile hybrids, the rates of gene exchange between species are low.

evolutionary biology

Spatial and temporal distribution of genome divergence among California populations of Aedes aegypti.

In the summer of 2013, Aedes aegypti Linnaeus was first detected in three cities in central California (Clovis, Madera and Menlo Park). It has now been detected in multiple locations in central and southern CA as far south as San Diego and Imperial Counties. A number of published reports suggest that CA populations have been established from multiple independent introductions. Here we report the first population genomics analyses of Ae. aegypti based on individual, field collected whole genome sequences. We analyzed 46 Ae. aegypti genomes to establish genetic relationships among populations from sites in California, Florida and South Africa. We identified 3 major genetic clusters within California; one that includes all sample sites in the southern part of the state (South of Tehachapi mountain range) plus the town of Exeter in central California and two additional clusters in central California. A lack of concordance between mitochondrial and nuclear genealogies suggests that the three founding populations were polymorphic for two main mitochondrial haplotypes prior to being introduced to California. One of these has been lost in the Clovis populations, possibly by a founder effect. Genome-wide comparisons indicate extensive differentiation between genetic clusters. Our observations support recent introductions of Ae. aegypti into California from multiple, genetically diverged source populations. Our data reveal signs of hybridization among diverged populations within CA. Genetic markers identified in this study will be of great value in pursuing classical population genetic studies which require larger sample sizes.

genomics

Predictors of intestinal inflammation in asymptomatic first-degree relatives of patients with Crohn’s disease

ObjectiveRelatives of individuals with Crohns disease (CD) carry an increased number of CD-associated genetic variants and are at increased risk of developing the disease. Multiple environmental and genetic factors contribute to this increased risk. We aimed to estimate the utility of genotype, smoking, family history, and a panel of biomarkers to predict risk in asymptomatic first-degree relatives (FDRs) of CD patients.\n\nDesignWe calculated a combined genotype (72 CD-associated genetic markers) and smoking relative risk score in 454 FDRs, and performed capsule endoscopy and collected 22 biomarkers in individuals from the highest and lowest risk quartiles. We then predicted small intestinal inflammation using genetic risk score, smoking status, number of relatives with CD, capsule transit time, and the panel of biomarkers in 124 individuals with complete data. Our principal analysis was to calculate the predictive utility from two machine learning classifiers: an elastic net and a random forest.\n\nResultsBoth classifiers successfully predicted FDRs with intestinal inflammation: elastic net (AUC=0.80, 95% CI: 0.62-0.98), random forest (AUC=0.87, 95% CI: 0.75-1.00). The elastic net selected a 3-predictor solution: CD family history (OR=1.31), genetic risk score (OR=1.14), and faecal calprotectin (OR=1.04). The same 3 variables were among the top 5 most important predictors as ranked by the random forest.\n\nConclusionA readily collectable panel of genetic risk variants, added to family history and faecal calprotectin, predicts those at greatest risk for developing CD with a good degree of accuracy.

epidemiology

The genomics of local adaptation in trees: Are we out of the woods yet?

There is substantial interest in uncovering the genetic basis of the traits underlying adaptive responses in tree species, as this information will ultimately aid conservation and industrial endeavors across populations, generations, and environments. Fundamentally, the characterization of such genetic bases is within the context of a genetic architecture, which describes the mutlidimensional relationship between genotype and phenotype through the identification of causative variants, their relative location within a genome, expression, pleiotropic effect, environmental influence, and degree of dominance, epistasis, and additivity. Here, we review theory related to polygenic local adaptation and contextualize these expectations with methods often used to uncover the genetic basis of traits important to tree conservation and industry. A broad literature survey suggests that most tree traits generally exhibit considerable heritability, that underlying quantitative genetic variation (QST) is structured more so across populations than neutral expectations (FST) in 69% of comparisons across the literature, and that single-locus associations often exhibit small estimated per-locus effects. Together, these results suggest differential selection across populations often acts on tree phenotypes underlain by polygenic architectures consisting of numerous small to moderate effect loci. Using this synthesis, we highlight the limits of using solely single-locus approaches to describe underlying genetic architectures and close by addressing hurdles and promising alternatives towards such goals, remark upon the current state of tree genomics, and identify future directions for this field. Importantly, we argue, the success of future endeavors should not be predicated on the shortcomings of past studies and will instead be dependent upon the application of theory to empiricism, standardized reporting, centralized open-access databases, and continual input and review of the communitys research.

evolutionary biology

A population phylogenetic view of mitochondrial heteroplasmy

The mitochondrion has recently emerged as an active player in a myriad of cellular processes. Additionally, it was recently shown that more than 200 diseases are known to be linked to variants in mitochondrial DNA or in nuclear genes interacting with mitochondria. This has reinvigorated interest in its biology and population genetics. Mitochondrial heteroplasmy, or genotypic variation of mitochondria within an individual, is now understood to be common in humans and important in human health. However, it is still not possible to make quantitative predictions about the inheritance of heteroplasmy and its proliferation within the body, partly due to the lack of an appropriate model. Here, we present a population-genetic framework for modeling mitochondrial heteroplasmy as a process that occurs on an ontogenetic phylogeny, with genetic drift and mutation changing heteroplasmy frequencies during the various developmental processes represented in the phylogeny. Using this framework, we develop a Bayesian inference method for inferring rates of mitochondrial genetic drift and mutation at different stages of human life. Applying the method to previously published heteroplasmy frequency data, we demonstrate a severe effective germline bottleneck comprised of the cumulative genetic drift occurring between the divergence of germline and somatic cells in the mother and the separation of germ layers in the offspring. Additionally, we find that the two somatic tissues we analyze here undergo tissue-specific bottlenecks during embryogenesis, less severe than the effective germline bottleneck, and that these somatic tissues experience little additional genetic drift during adulthood. We conclude with a discussion of possible extensions of the ontogenetic phylogeny framework and its possible applications to other ontogenetic processes in addition to mitochondrial heteroplasmy.

genomics

The radiation of alopiine clausiliids in the Sicilian Channel (Central Mediterranean): phylogeny, patterns of morphological diversification and implications for taxonomy and conservation of Muticaria and Lampedusa

The phylogeny, biogeography and taxonomy of the alopiine clausiliids of the Sicilian Channel, belonging to the genera Lampedusa and Muticaria, were investigated using morphological (shell characters and anatomy of the reproductive system) and genetic (sequencing of a fragment of the mitochondrial large ribosomal subunit 16S rRNA, and the nuclear internal transcriber spacer 1, ITS-1 rRNA) data. Classically, the genus Lampedusa includes three species: L. imitatrix and L. melitensis occurring in circumscribed localities in western Malta and on the islet of Filfla, and L. lopadusae on Lampedusa and Lampione. The genus Muticaria includes two species in southeastern Sicily (M. siracusana and M. neuteboomi) and one in the Maltese islands (M. macrostoma), which is usually subdivided into four entities based on shell characters (macrostoma on Gozo, Comino, Cominotto and central-eastern Malta; mamotica in southeastern Gozo; oscitans on Gozo and central-western Malta; scalaris in northwestern Malta). These have sometimes been considered as subspecies and sometimes as mere morphs.\n\nThe Lampedusa of Lampedusa and Lampione form a well distinct clade from those of the Maltese Islands. The population of Lampione islet is a genetically distinct geographic form that deserves formal taxonomic recognition (as L. nodulosa or L. l. nodulosa). The Lampedusa of Malta are morphologically distinct evolutionary lineages with high levels of genetic divergence and are confirmed as distinct species (L. imitatrix and L. melitensis).\n\nThe Muticaria constitute a clearly different monophyletic clade divided into three geographical lineages corresponding to the Sicilian, Maltese and Gozitan populations. The Sicilian Muticaria form two morphologically and genetically distinguishable subclades that may either be considered subspecies of a polytypic species, or two distinct species. The relationships of Maltese and Gozitan Muticaria are complex. Two of the three Maltese morphotypes resulted monophyletic (oscitans and scalaris) while the other was separated in two lineages (macrostoma); however this picture may be biased as only few samples of macrostoma were available to study. The Gozitan morphotypes (macrostoma, mamotica and oscitans) where resolved as polyphyletic but with clear molecular evidence of mixing in some cases, indicating possible relatively recent differentiation of the Gozitan Muticaria or repetitive secondary contacts between different morphotypes. Definitive taxonomic conclusions from these results are premature. Maltese Muticaria could be subdivided into three taxa according to morphological and molecular data (M. macrostoma or M. m. macrostoma, M. oscitans or M. m. oscitans and M. scalaris or M. m scalaris). Gozitan Muticaria could be considered a distinct polytypic species (for which the oldest available name is Muticaria mamotica) subdivided into subspecies showing a morphological range from macrostoma-like to mamotica-like and oscitans like.\n\nOnly the two Maltese species of Lampedusa are legally protected (by the European Unions Habitats Directive and Maltese national legislation). The present study has shown that the alopiine clausiliids of the Sicilian Channel constitute a number of genetically and/or morphologically distinct populations that represent important pools of genetic diversity, with, in some cases, a very circumscribed distribution. As such, these populations deserve legal protection and management. It is argued that without formal taxonomic designation, it would be difficult to extend international legal protection to some of the more threatened of these populations.

zoology

Division of labor during biofilm matrix production

Organisms as simple as bacteria can engage in complex collective actions, such as group motility and fruiting body formation. Some of these actions involve a division of labor, where phenotypically specialized clonal subpopulations, or genetically distinct lineages cooperate with each other by performing complementary tasks. Here, we combine experimental and computational approaches to investigate potential benefits arising from division of labor during biofilm matrix production. We show that both phenotypic and genetic strategies for a division of labor can promote collective biofilm formation in the soil bacterium Bacillus subtilis. In this species, biofilm matrix consists of two major components; EPS and TasA. We observed that clonal groups of B. subtilis phenotypically segregate into three subpopulations composed of matrix non-producers, EPS-producers, and generalists, which produce both EPS and TasA. This incomplete phenotypic specialization was outperformed by a genetic division of labor, where two mutants, engineered as specialists, complemented each other by exchanging EPS and TasA. The relative fitness of the two mutants displayed a negative frequency dependence both in vitro and on plant roots, with strain frequency reaching a stable equilibrium at 30% TasA-producers, corresponding exactly to the population composition where group productivity is maximized. Using individual-based modelling, we show that asymmetries in strain ratio can arise due to differences in the relative benefits that matrix compounds generate for the collective; and that genetic division of labor can be favored when it breaks metabolic constraints associated with the simultaneous production of two matrix components.\n\nHighlights- matrix components EPS and TasA are costly public goods in B. subtilis biofilms\n\n- genetic division of labor using {Delta}eps and {Delta}tasA fosters maximal biofilm productivity\n\n- {Delta}eps and {Delta}tasA cooperation is evolutionary stable in laboratory and ecological systems\n\n- costly metabolic coupling of public goods favors genetic division of labor

microbiology

Genome-wide association across Saccharomyces cerevisiae strains reveals substantial variation in underlying gene requirements for toxin tolerance.

Cellulosic plant biomass is a promising sustainable resource for generating alternative biofuels and biochemicals with microbial factories. But a remaining bottleneck is engineering microbes that are tolerant of toxins generated during biomass processing, because mechanisms of toxin defense are only beginning to emerge. Here, we exploited natural diversity in 165 Saccharomyces cerevisiae strains isolated from diverse geographical and ecological niches, to identify mechanisms of hydrolysate-toxin tolerance. We performed genome-wide association (GWA) analysis to identify genetic variants underlying toxin tolerance, and gene knockouts and allele-swap experiments to validate the involvement of implicated genes. In the process of this work, we uncovered a surprising difference in genetic architecture depending on strain background: in all but one case, knockout of implicated genes had a significant effect on toxin tolerance in one strain, but no significant effect in another strain. In fact, whether or not the gene was involved in tolerance in each strain background had a bigger contribution to strain-specific variation than allelic differences. Our results suggest a major difference in the underlying network of causal genes in different strains, suggesting that mechanisms of hydrolysate tolerance are very dependent on the genetic background. These results could have significant implications for interpreting GWA results and raise important considerations for engineering strategies for industrial strain improvement.\n\nAuthor summaryUnderstanding the genetic architecture of complex traits is important for elucidating the genotype-phenotype relationship. Many studies have sought genetic variants that underlie phenotypic variation across individuals, both to implicate causal variants and to inform on architecture. Here we used genome-wide association analysis to identify genes and processes involved in tolerance of toxins found in plant-biomass hydrolysate, an important substrate for sustainable biofuel production. We found substantial variation in whether or not individual genes were important for tolerance across genetic backgrounds. Whether or not a gene was important in a given strain background explained more variation than the alleleic differences in the gene. These results suggest substantial variation in gene contributions, and perhaps underlying mechanisms, of toxin tolerance.

genomics

Strong hybrid male incompatibilities impede the spread of a selfish chromosome between populations of a fly.

Meiotically driving sex chromosomes manipulate gametogenesis to increase their transmission at a cost to the rest of the genome. The intragenomic conflicts they produce have major impacts on the ecology and evolution of their host species. However their ecological dynamics remain poorly understood. Simple population genetic models predict meiotic drivers will rapidly reach fixation in a population and spread across a landscape. In contrast, natural populations commonly show spatial variation in the frequency of drivers, with drive present in clines or mosaics across species ranges. For example, Drosophila subobscura harbours a Sex Ratio distorting drive chromosome (\"SRs\") at 15-25% frequency in North Africa, present at less than 2% frequency in adjacent Southern Spain and absent in other European populations. Here, we investigate the forces preventing the spread of the driver northward. We show that SRs has remained at a constant frequency in North Africa, and failed to spread in Spain. We find strong evidence in favour of our first hypothesis, genetic incompatibility between SRs and Spanish autosomal background. When we cross SRs from North Africa onto Spanish genetic backgrounds we observe strong SRs specific incompatibilities in hybrids. The incompatibilities increase in severity in F2 male hybrids, leading to almost complete infertility. We find no evidence supporting a second hypothesis, that there is resistance to drive in Spanish populations. We conclude that the source of the stepped frequency variation is genetic incompatibility between the SRs chromosome and the genetic backgrounds of the adjacent population, preventing SRs spreading northward. The low frequency of SRs in South Spain is consistent with recurrent gene flow across the Strait of Gibraltar combined with selection against the SRs element through genetic incompatibility. This demonstrates that incompatibilities between drive chromosomes and naive populations can prevent the spread of drive between populations, at a continental scale.

evolutionary biology

A high-quality sequence of Rosa chinensis to elucidate genome structure and ornamental traits

Rose is the worlds most important ornamental plant with economic, cultural and symbolic value. Roses are cultivated worldwide and sold as garden roses, cut flowers and potted plants. Rose has a complex genome with high heterozygosity and various ploidy levels. Our objectives were (i) to develop the first high-quality reference genome sequence for the genus Rosa by sequencing a doubled haploid, combining long and short read sequencing, and anchoring to a high-density genetic map and (ii) to study the genome structure and the genetic basis of major ornamental traits.\n\nWe produced a haploid rose line from R. chinensis Old Blush and generated the first rose genome sequence at the pseudo-molecule scale (512 Mbp with N50 of 3.4 Mb and L75 of 97). The sequence was validated using high-density diploid and tetraploid genetic maps. We delineated hallmark chromosomal features including the pericentromeric regions through annotation of TE families and positioned centromeric repeats using FISH. Genetic diversity was analysed by resequencing eight Rosa species. Combining genetic and genomic approaches, we identified potential genetic regulators of key ornamental traits, including prickle density and number of flower petals. A rose APETALA2 homologue is proposed to be the major regulator of petals number in rose. This reference sequence is an important resource for studying polyploidisation, meiosis and developmental processes as we demonstrated for flower and prickle development. This reference sequence will also accelerate breeding through the development of molecular markers linked to traits, the identification of the genes underlying them and the exploitation of synteny across Rosaceae.

genomics

Bayesian phylodynamic inference with complex models

Population genetic modeling can enhance Bayesian phylogenetic inference by providing a realistic prior on the distribution of branch lengths and times of common ancestry.The parameters of a population genetic model may also have intrinsic importance, and simultaneous estimation of a phylogeny and model parameters has enabled phylodynamic inference of population growth rates, reproduction numbers, and effective population size through time. Phylodynamic inference based on pathogen genetic sequence data has emerged as useful supplement to epidemic surveillance, however commonly-used mechanistic models that are typically fitted to non-genetic surveillance data are rarely fitted to pathogen genetic data due to a dearth of software tools, and the theory required to conduct such inference has been developed only recently. We present a framework for coalescent-based phylogenetic and phylodynamic inference which enables highly-flexible modeling of demographic and epidemiological processes. This approach builds upon previous structured coalescent approaches and includes enhancements for computational speed, accuracy, and stability. A flexible markup language is described for translating parametric demographic or epidemiological models into a structured coalescent model enabling simultaneous estimation of demographic or epidemiological parameters and time-scaled phylogenies. We demonstrate the utility of these approaches by fitting compartmental epidemiological models to Ebola virus and Influenza A virus sequence data, demonstrating how important features of these epidemics, such as the reproduction number and epidemic curves, can be gleaned from genetic data. These approaches are provided as an open-source package PhyDyn for the BEAST phylogenetics platform.

bioinformatics

Principal component based adaptive association test of multiple traits using GWAS summary statistics

Genetics hold great promise to precision medicine by tailoring treatment to the individual patient based on their genetic profiles. Toward this goal, many large-scale genome-wide association studies (GWAS) have been performed in the last decade to identify genetic variants associated with various traits and diseases. They have successfully identified tens of thousands of disease-related variants. However they have explained only a small proportion of the overall trait heritability for most traits and are of very limited clinical use. This is partly owing to the small effect sizes of most genetic variants, and the common practice of \"testing association between one trait and one genetic variant at a time\" in most GWAS, even when multiple related traits are often measured for each individual. Increasing evidence suggests that many genetic variants can influence multiple traits simultaneously, and we can gain more power by testing association of multiple traits simultaneously. It is appealing to develop novel multi-trait association test methods that need only GWAS summary data, since it is generally very hard to access the individual-level GWAS phenotype and genotype data.\n\nMost existing GWAS summary data based association test methods have relied on ad hoc approach or crude Monte Carlo approximation. In this paper we develop rigorous statistical methods for efficient and powerful multi-trait association test. We develop robust and efficient methods to accurately estimate the marginal trait correlation matrix using only GWAS summary data. We construct the principal component (PC) based association test from the summary statistics. PC based test has optimal power when the underlying multi-trait signal can be captured by the first PC, and otherwise it will have suboptimal performance. We develop an adaptive test by optimally weighting the PC based test and the omnibus chi-square test to achieve robust performance under various scenarios. We develop efficient numerical algorithms to compute the analytical p-values for all the proposed tests without the need of Monte Carlo sampling. We illustrate the utility of proposed methods through application to the GWAS meta-analysis summary data for multiple lipids and glycemic traits. We identify multiple novel loci that were missed by individual trait based association test.\n\nAll the proposed methods are implemented in an R package available at http://www.github.com/baolinwu/MTAR. The developed R programs are extremely efficient: it takes less than two minutes to compute the list of genome-wide significant SNPs for all proposed multi-trait tests for the lipids GWAS summary data with 2.5 million SNPs on a single Linux desktop.

bioinformatics