Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Reference trait analysis reveals correlations between gene expression and quantitative traits in disjoint samples

Systems genetic analysis of complex traits involves the integrated analysis of genetic, genomic, and disease related measures. However, these data are often collected separately across multiple study populations, rendering direct correlation of molecular features to complex traits impossible. Recent transcriptome-wide association studies (TWAS) have harnessed gene expression quantitative trait loci (eQTL) to associate unmeasured gene expression with a complex trait in genotyped individuals, but this approach relies primarily on strong eQTLs. We propose a simple and powerful alternative strategy for correlating independently obtained sets of complex traits and molecular features. In contrast to TWAS, our approach gains precision by correlating complex traits through a common set of continuous phenotypes instead of genetic predictors, and can identify transcript-trait correlations for which the regulation is not genetic. In our approach, a set of multiple quantitative \"reference\" traits is measured across all individuals, while measures of the complex trait of interest and transcriptional profiles are obtained in disjoint sub-samples. A conventional multivariate statistical method, canonical correlation analysis, is used to relate the reference traits and traits of interest in order to identify gene expression correlates. We evaluate power and sample size requirements of this methodology, as well as performance relative to other methods, via extensive simulation and analysis of a behavioral genetics experiment in 258 Diversity Outbred mice involving two independent sets of anxiety-related behaviors and hippocampal gene expression. After splitting the dataset and hiding one set of anxiety-related traits in half the samples, we identified transcripts correlated with the hidden traits using the other set of anxiety-related traits and exploiting the highest canonical correlation (R = 0.69) between the trait datasets. We demonstrate that this approach outperforms TWAS in identifying associated transcripts. Together, these results demonstrate the validity, reliability, and power of the reference trait method for identifying relations between complex traits and their molecular substrates.\n\nAUTHOR SUMMARYSystems genetics exploits natural genetic variation and high-throughput measurements of molecular intermediates to dissect genetic contributions to complex traits. An important goal of this strategy is to correlate molecular features, such as transcript or protein abundance, with complex traits. For practical, technical, or financial reasons, it may be impossible to measure complex traits and molecular intermediates on the same individuals. Instead, in some cases these two sets of traits may be measured on independent cohorts. We outline a method, reference trait analysis, for identifying molecular correlates of complex traits in this scenario. We show that our method powerfully identifies complex trait correlates across a wide range of parameters that are biologically plausible and experimentally practical. Furthermore, we show that reference trait analysis can identify transcripts correlated to a complex trait more accurately than approaches such as TWAS that use genetic variation to predict gene expression. Reference trait analysis will contribute to furthering our understanding of variation in complex traits by identifying molecular correlates of complex traits that are measured in different individuals.

genomics

Spatial localization of recent ancestors for admixed individuals

Ancestry analysis from genetic data plays a critical role in studies of human disease and evolution. Recent work has introduced explicit models for the geographic distribution of genetic variation and has shown that such explicit models yield superior accuracy in ancestry inference over non-model-based methods. Here we extend such work to introduce a method that models admixture between ancestors from multiple sources across a geographic continuum. We devise efficient algorithms based on hidden Markov models to localize on a map the recent ancestors (e.g. grandparents) of admixed individuals, joint with assigning ancestry at each locus in the genome. We validate our methods using empirical data from individuals with mixed European ancestry from the POPRES study and show that our approach is able to localize their recent ancestors within an average of 470Km of the reported locations of their grandparents. Furthermore, simulations from real POPRES genotype data show that our method attains high accuracy in localizing recent ancestors of admixed individuals in Europe (an average of 550Km from their true location for localization of 2 ancestries in Europe, 4 generations ago). We explore the limits of ancestry localization under our approach and find that performance decreases as the number of distinct ancestries and generations since admixture increases. Finally, we build a map of expected localization accuracy across admixed individuals according to the location of origin within Europe of their ancestors.\n\nAuthor SummaryInferring ancestry from genetic data forms a fundamental problem with applications ranging from localizing disease genes to inference of human history. Recent approaches have introduced models of genetic variation as a function of geography and have shown that such models yield high accuracies in ancestry inference from genetic data. In this work we propose methods for modeling the mixing of genetic data from different sources (i.e. admixture process) in a genetic-geographic continuum and show that using these methods we can accurately infer the ancestry of the recent ancestors (e.g. grandparents) from genetic data.

Genetics

Robust Population Structure Inference and Correction in the Presence of Known or Cryptic Relatedness

Population structure inference with genetic data has been motivated by a variety of applications in population genetics and genetic association studies. Several approaches have been proposed for the identification of genetic ancestry differences in samples where study participants are assumed to be unrelated, including principal components analysis (PCA), multi-dimensional scaling (MDS), and model-based methods for proportional ancestry estimation. Many genetic studies, however, include individuals with some degree of relatedness, and existing methods for inferring genetic ancestry fail in related samples. We present a method, PC-AiR, for robust population structure inference in the presence of known or cryptic relatedness. PC-AiR utilizes genome-screen data and an efficient algorithm to identify a diverse subset of unrelated individuals that is representative of all ancestries in the sample. The PC-AiR method directly performs PCA on the identified ancestry representative subset and then predicts components of variation for all remaining individuals based on genetic similarities. In simulation studies and in applications to real data from Phase III of the HapMap Project, we demonstrate that PC-AiR provides a substantial improvement over existing approaches for population structure inference in related samples. We also demonstrate significant efficiency gains, where a single axis of variation from PC-AiR provides better prediction of ancestry in a variety of structure settings than using ten (or more) components of variation from widely used PCA and MDS approaches. Finally, we illustrate that PC-AiR can provide improved population stratification correction over existing methods in genetic association studies with population structure and relatedness.

Genetics

Do gametes woo? Evidence for non-random unions at fertilization

A fundamental tenet of inheritance in sexually reproducing organisms such as humans and laboratory mice is that genetic variants combine randomly at fertilization, thereby ensuring a balanced and statistically predictable representation of inherited variants in each generation. This principle is encapsulated in Mendels First Law. But exceptions are known. With transmission ratio distortion (TRD), particular alleles are preferentially transmitted to offspring without reducing reproductive productivity. Preferential transmission usually occurs in one sex but not both and is not known to require interactions between gametes at fertilization. We recently discovered, in our work in mice and in other reports in the literature, instances where any of 12 mutant genes bias fertilization, with either too many or too few heterozygotes and too few homozygotes, depending on the mutant gene and on dietary conditions. Although such deviations are usually attributed to embryonic lethality of the under-represented genotypes, the evidence is more consistent with genetically-determined preferences for specific combinations of egg and sperm at fertilization that results in genotype bias without embryo loss. These genes and diets could bias fertilization in at least three not mutually exclusive ways. They could trigger a reversal in the order of meiotic divisions during oogenesis so that the genetics of fertilizing sperm elicits preferential chromatid segregation, thereby dictating which allele remains in the egg versus the 2nd polar body. Bias could also result from genetic- and diet-induced anomalies in polyamine metabolism on which function of haploid gametes normally depends. Finally, secreted and cell-surface factors in female reproductive organs could control access of sperm to eggs based on their genetic content. This unexpected discovery of genetically-biased fertilization in mice could yield insights about the molecular and cellular interactions between sperm and egg at fertilization, with implications for our understanding of inheritance, reproduction, population genetics, and medical genetics.

genetics

Improving on a modal-based estimation method: model averaging for consistent and efficient estimation in Mendelian randomization when a plurality of candidate instruments are valid

BackgroundA robust method for Mendelian randomization does not require all genetic variants to be valid instruments to give consistent estimates of a causal parameter. Several such methods have been developed, including a mode-based estimation method giving consistent estimates if a plurality of genetic variants are valid instruments; that is, there is no larger subset of invalid instruments estimating the same causal parameter than the subset of valid instruments.\n\nMethodsWe here develop a model averaging method that gives consistent estimates under the same plurality of valid instruments assumption. The method considers a mixture distribution of estimates derived from each subset of genetic variants. The estimates are weighted such that subsets with more genetic variants receive more weight, unless variants in the subset have heterogeneous causal estimates, in which case that subset is severely downweighted. The mode of this mixture distribution is the causal estimate. This heterogeneity-penalized model averaging method has several technical advantages over the previously proposed mode-based estimation method.\n\nResultsThe heterogeneity-penalized model averaging method outperformed the mode-based estimation in terms of effciency and outperformed other robust methods in terms of Type 1 error rate in an extensive simulation analysis. The proposed method suggests two distinct mechanisms by which inflammation affects coronary heart disease risk, with subsets of variants suggesting both positive and negative causal effects.\n\nConclusionsThe heterogeneity-penalized model averaging method is an additional robust method for Mendelian randomization with excellent theoretical and practical properties, and can reveal features in the data such as the presence of multiple causal mechanisms. (249 words)\n\nKey messagesO_LIWe propose a heterogeneity-penalized model averaging method that gives consistent causal estimates if a weighted plurality of the genetic variants are valid instruments.\nC_LIO_LIThe method calculates causal estimates based on all subsets of genetic variants, and upweights subsets containing several genetic variants with similar causal estimates.\nC_LIO_LIThe method is asymptotically effcient and does not rely on bootstrapping to obtain a confidence interval, nor is the confidence interval constrained to be symmetric.\nC_LIO_LIIn particular, the confidence interval can include multiple disjoint intervals, suggesting the presence of multiple causal mechanisms by which the risk factor influences the outcome.\nC_LIO_LIThe method can incorporate biological knowledge to upweight the contribution of genetic variants with stronger plausibility of being valid instruments.\nC_LI

genetics

Functional testing of a human PBX3 variant in zebrafish reveals a potential modifier role in congenital heart defects

Whole-genome and whole-exome sequencing efforts are increasingly identifying candidate genetic variants associated with human disease. However, predicting and testing the pathogenicity of a genetic variant remains challenging. Genome editing allows for the rigorous functional testing of human genetic variants in animal models. Congenital heart defects (CHDs) are a prominent example of a human disorder with complex genetics. An inherited sequence variant in the human PBX3 gene (PBX3 p.A136V) has previously been shown to be enriched in a CHD patient cohort, indicating that the PBX3 p.A136V variant could be a modifier allele for CHDs. PBX genes encode TALE (Three Amino acid Loop Extension)-class homeodomain-containing DNA-binding proteins with diverse roles in development and disease and are required for heart development in mouse and zebrafish. Here we use CRISPR-Cas9 genome editing to directly test whether this PBX gene variant acts as a genetic modifier in zebrafish heart development. We used a single-stranded oligodeoxynucleotide to precisely introduce the human PBX3 p.A136V variant in the homologous zebrafish pbx4 gene (pbx4 p.A131V). We find that zebrafish that are homozygous for pbx4 p.A131V are viable as adults. However, we show that the pbx4 p.A131V variant enhances the embryonic cardiac morphogenesis phenotype caused by loss of the known cardiac specification factor, Hand2. Our study is the first example of using precision genome editing in zebrafish to demonstrate a function for a human disease-associated single nucleotide variant of unknown significance. Our work underscores the importance of testing the roles of inherited variants, not just de novo variants, as genetic modifiers of CHDs. Our study provides a novel approach toward advancing our understanding of the complex genetics of CHDs.\n\nSummary statementOur study provides a novel example of using genome editing in zebrafish to demonstrate how a human DNA sequence variant of unknown significance may contribute to the complex genetics of congenital heart defects.

genetics

Exploring Various Polygenic Risk Scores for Basal Cell Carcinoma, Cutaneous Squamous Cell Carcinoma and Melanoma in the Phenomes of the Michigan Genomics Initiative and the UK Biobank

Polygenic risk scores (PRS) are designed to serve as a single summary measure, condensing information from a large number of genetic variants associated with a disease. They have been used for stratification and prediction of disease risk. The construction of a PRS often depends on the purpose of the study, the available data/summary estimates, and the underlying genetic architecture of a disease. In this paper, we consider several choices for constructing a PRS using summary data obtained from various publicly-available sources including the UK Biobank and evaluate their abilities to predict outcomes derived from electronic health records (EHR). Weexamine the three most common skin cancer subtypes in the USA: basal cellcarcinoma, cutaneous squamous cell carcinoma, and melanoma. The genetic risk profiles of subtypes may consist of both shared and unique elements and we construct PRS to understand the common versus distinct etiology. This study is conducted using data from 30,702 unrelated, genotyped patients of recent European descent from the Michigan Genomics Initiative (MGI), a longitudinal biorepository effort within Michigan Medicine. Using these PRS for various skin cancer subtypes, we conduct a phenome-wide association study (PheWAS) within the MGI data to evaluate their association with secondary traits. PheWAS results are then replicated using population-based UK Biobank data. We develop an accompanying visual catalog called PRSweb that provides detailed PheWAS results and allows users to directly compare different PRS construction methods. The results of this study can provide guidance regarding PRS construction in future PRS-PheWAS studies using EHR data involving disease subtypes.\n\nAuthor summaryIn the study of genetically complex diseases, polygenic risk scores synthesize information from multiple genetic risk factors to provide insight into a patients risk of developing a disease based on his/her genetic profile. These risk scores can be explored in conjunction with health and disease information available in the electronic medical records. They may be associated with diseases that may be related to or precursors of the underlying disease of interest. Limited work is available guiding risk score construction when the goal is to identify associations across the medical phenome. In this paper, we compare different polygenic risk score construction methods in terms of their relationships with the medical phenome. We further propose methods for using these risk scores to decouple the shared and unique genetic profiles of related diseases and to explore related diseases shared and unique secondary associations. Leveraging and harnessing the rich data resources of the Michigan Genomics Initiative, a biorepository effort at Michigan Medicine, and the larger population-based UK Biobank study, we investigated the performance of genetic risk profiling methods for the three most common types of skin cancer: melanoma, basal cell carcinoma and squamous cell carcinoma.

genetics

Cleft lip/palate and educational attainment: cause, consequence, or correlation? A Mendelian randomization study

ImportancePrevious studies have found that children born with a non-syndromic form of cleft lip and/or palate have lower-than-average educational attainment. These differences could be due to a genetic predisposition to low intelligence and academic performance, factors arising due to the cleft phenotype (such as school absence, social stigmatization and impaired speech and language development), or confounding by the prenatal environment. A clearer understanding of this mechanism will inform development of interventions to improve educational attainment in individuals born with a cleft, which could have wide-ranging knock-on effects on their quality of life.\n\nObjectiveTo assess evidence for the hypothesis that common variant genetic liability to non-syndromic cleft lip with or without cleft palate (nsCL/P) influences educational attainment.\n\nDesignUsing summary data from genome-wide association studies (GWAS), we performed Linkage Disequilibrium (LD)-score regression and two-sample Mendelian randomization to evaluate the relationship between genetic liability to nsCL/P (GWAS n=3,987) and educational attainment (GWAS n=766,345), and intelligence (GWAS n=257,828).\n\nResultsThere was little evidence for shared genetic aetiology between nsCL/P and educational attainment (rg -0.03, 95% CI -0.14 to 0.08, P 0.58; {beta}MR 0.002, 95% CI -0.001 to 0.005, P 0.417) or intelligence (rg -0.01, 95% CI -0.12 to 0.10, P 0.85; {beta}MR 0.002, 95% CI -0.010 to 0.014, P 0.669).\n\nConclusions and relevanceCommon genetic variants are unlikely to predispose individuals born with nsCL/P to low educational attainment or intelligence. This information will help tailor clinical-, school-, social- and family-level interventions to improve educational attainment in this group.\n\nKey PointsO_ST_ABSQuestionC_ST_ABSDo children born with a non-syndromic cleft lip with or without palate (nsCL/P) have lower-than average academic achievement because of an underlying genetic predisposition to educational attainment and/or intelligence?\n\nFindingsThere was little evidence for shared common variant genetic correlation between nsCL/P, educational attainment and intelligence.\n\nMeaningCommon genetic variants are unlikely to predispose individuals born with nsCL/P to low educational attainment or intelligence. This information will help tailor clinical-, school-, social- and family-level interventions to improve educational attainment in this group.

genetics

Using DNA from mothers and children to study parental investment in children’s educational attainment

This study tested implications of new genetic discoveries for understanding the association between parental investment and childrens educational attainment. A novel design matched genetic data from 860 British mothers and their children with home-visit measures of parenting: the E-Risk Study. Three findings emerged. First, both mothers and childrens education-associated genetics, summarized in a genome-wide polygenic score, predicted parenting -- a gene-environment correlation. Second, accounting for genetic influences slightly reduced associations between parenting and childrens attainment -- indicating some genetic confounding. Third, mothers genetics influenced childrens attainment over and above genetic mother-to-child transmission, via cognitively-stimulating parenting -- an environmentally-mediated effect. Findings imply that, when interpreting parents effects on children, environmentalists must consider genetic transmission, but geneticists must also consider environmental transmission.

genetics

Tuberculosis susceptibility and vaccine protection are independently controlled by host genotype

The outcome of Mycobacterium tuberculosis (Mtb) infection and the immunological response to the Bacille Calmette Guerin (BCG) vaccine are highly variable in humans. Deciphering the relative importance of host genetics, environment, and vaccine preparation on BCG efficacy has proven difficult in natural populations. We developed a model system that captures the breadth of immunological responses observed in outbred individuals, which can be used to understand the contribution of host genetics to vaccine efficacy. This system employs a panel of highly-diverse inbred mouse strains, consisting of the founders and recombinant progeny of the \"Collaborative Cross\". Unlike natural populations, the structure of this panel allows the serial evaluation of genetically-identical individuals and quantification of genotype-specific effects of interventions such as vaccination. When analyzed in the aggregate, our panel resembled natural populations in several important respects; the animals displayed a broad range of Mtb susceptibility, varied in their immunological response to infection, and were not durably protected by BCG vaccination. However, when analyzed at the genotype level, we found that these phenotypic differences were heritable. Mtb susceptibility varied between lines, from extreme sensitivity to progressive Mtb clearance. Similarly, only a minority of the genotypes was protected by vaccination. BCG efficacy was genetically separable from susceptibility, and the lack of efficacy in the aggregate analysis was driven by nonresponsive lines that mounted a qualitatively distinct response to infection. These observations support an important role for host genetic diversity in determining BCG efficacy, and provide a new resource to rationally develop more broadly efficacious vaccines.\n\nImportance: Tuberculosis (TB) remains an urgent global health crisis, and the efficacy of the currently used TB vaccine, M. bovis BCG, is highly variable. The design of more broadly-efficacious vaccines depends on understanding the factors that limit the protection imparted by BCG. While these complex factors are difficult to disentangle in natural populations, we used a model population of mice to understand the role of host genetic composition to BCG efficacy. We found that the ability of BCG to protect an individual genotype was remarkably variable. BCG efficacy did not depend on the intrinsic susceptibility of the animal, but instead correlated with qualitative differences in the immune response to the pathogen. These studies suggest that host genetic polymorphism is a critical determinant of vaccine efficacy and provides a model system to develop interventions that will be useful in genetically diverse populations.

Microbiology

Deconstructing isolation-by-distance: the genomic consequences of limited dispersal

Geographically limited dispersal can shape genetic population structure and result in a correlation between genetic and geographic distance, commonly called isolation-bydistance. Despite the prevalence of isolation-by-distance in nature, to date few studies have empirically demonstrated the processes that generate this pattern, largely because few populations have direct measures of individual dispersal and pedigree information. Intensive, long-term demographic studies and exhaustive genomic surveys in the Florida Scrub-Jay (Aphelocoma coerulescens) provide an excellent opportunity to investigate the influence of dispersal on genetic structure. Here, we used a panel of genome-wide SNPs and extensive pedigree information to explore the role of limited dispersal in shaping patterns of isolation-by-distance in both sexes, and at an exceedingly fine spatial scale (within ~10 km). Isolation-by-distance patterns were stronger in male-male and male-female comparisons than in female-female comparisons, consistent with observed differences in dispersal propensity between the sexes. Using the pedigree, we demonstrated how various genealogical relationships contribute to fine-scale isolation-by-distance. Simulations using field-observed distributions of male and female natal dispersal distances showed good agreement with the distribution of geographic distances between breeding individuals of different pedigree relationship classes. Furthermore, we extended Malecots theory of isolation-by-distance by building coalescent simulations parameterized by the observed dispersal curve, population density, and immigration rate, and showed how incorporating these extensions allows us to accurately reconstruct observed sex-specific isolation-by-distance patterns in autosomal and Z-linked SNPs. Therefore, patterns of fine-scale isolation-by-distance in the Florida Scrub-Jay can be well understood as a result of limited dispersal over contemporary timescales.\n\nAuthor SummaryDispersal is a fundamental component of the life history of most organisms and therefore influences many biological processes. Dispersal is particularly important in creating genetic structure on the landscape. We often observe a pattern of decreased genetic relatedness between individuals as geographic distances increases, or isolation-by-distance. This pattern is particularly pronounced in organisms with extremely short dispersal distances. Despite the ubiquity of isolation-by-distance patterns in nature, there are few examples that explicitly demonstrate how limited dispersal influences spatial genetic structure. Here we investigate the processes that result in spatial genetic structure using the Florida Scrub-Jay, a bird with extremely limited dispersal behavior and extensive genome-wide data. We take advantage of the long-term monitoring of a contiguous population of Florida Scrub-Jays, which has resulted in a detailed pedigree and measurements of dispersal for hundreds of individuals. We show how limited dispersal results in close genealogical relatives living closer together geographically, which generates a strong pattern of isolation-by-distance at an extremely small spatial scale (<10 km) in just a few generations. Given the detailed dispersal, pedigree, and genomic data, we can achieve a fairly complete understanding of how dispersal shapes patterns of genetic diversity over short spatial scales.

evolutionary biology

Network based conditional genome wide association analysis of human metabolomics

BackgroundGenome-wide association studies (GWAS) have identified hundreds of loci influencing complex human traits, however, their biological mechanism of action remains mostly unknown. Recent accumulation of functional genomics ( omics) including metabolomics data opens up opportunities to provide a new insight into the functional role of specific changes in the genome. Functional genomic data are characterized by high dimensionality, presence of (strong) statistical dependencies between traits, and, potentially, complex genetic control. Therefore, analysis of such data asks for development of specific statistical genetic methods.\n\nResultsWe propose a network-based, conditional approach to evaluate the impact of genetic variants on omics phenotypes (conditional GWAS, cGWAS). For each trait of interest, based on biological network, we select a set of other traits to be used as covariates in GWAS. The network could be reconstructed either from biological pathway databases or directly from the data. We evaluated our approach using data from a population-based KORA study (n=1,784, 1.7 M SNPs) with measured metabolomics data (151 metabolites) and demonstrated that our approach allows for identification of up to five additional loci not detected by conventional GWAS. We show that this gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.\n\nConclusionsWe demonstrate that in context of metabolomics network-based, conditional genome-wide association analysis is able to dramatically increase power of identification of loci with specific contra-intuitive pleiotropic architecture. Our method has modest computational costs, can utilize summary level GWAS data, and is applicable to other omics data types. We anticipate that application of our method to new and existing data sets will facilitate progress in understanding genetic bases of control of molecular and complex phenotypes.\n\nShort abstractWe propose a network-based, conditional approach for genome-wide analysis of multivariate omics phenotypes. Our methods can incorporate prior biological knowledge about biological pathways from external sources. We evaluated our approach using metabolomics data and demonstrated that our approach has bigger power and allows for identification of additional loci. We show that gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.

genomics

Mitochondrial genome variation affects the mutation rate of the nuclear genome in Drosophila melanogaster

Mutations are the raw material for evolutionary change. While the mutation rate has been thought constant between individuals, recent research has shown that poor genetic condition can elevate the mutation rate. Mitonuclear genetic conflict is a potential source of poor genetic condition, and considering the high mutation rate of mitochondrial genomes, there should be ample scope for mitochondrial mutations to interfere with genetic condition, with concomitant effects on the nuclear mutation rate. Moreover, because theory suggests mitochondrial genetic effects will often be male-biased, such effects could be more strongly felt in males than females. Here, by mating irradiated male Drosophila melanogaster to isogenic females bearing six distinct mitochondrial haplotypes, we tested whether mitochondrial genetic variation affects DNA repair capacity, and whether effects of mutation load on reproductive function are shaped by interactions between sex and mitochondrial haplotype. We found mitochondrial genetic effects on DNA repair, and that the mutational variance of reproductive fitness was higher in males bearing haplotypes characterized by high female fitness. These results suggest that mitochondrial genome variation may affect the mutation rate, and that induced mutations interact more strongly with male than female reproductive function. The potential for haplotype-specific effects on the nuclear mutation rate has broad implications for evolutionary dynamics, such as the accumulation of genetic load, adaptive potential, and the evolution of sexual dimorphism.

evolutionary biology

A Genome-wide Association and Admixture Mapping Study of Bronchodilator Drug Response in African Americans with Asthma

BackgroundShort-acting B2-adrenergic receptor agonists (SABAs) are the most commonly prescribed asthma medications worldwide. Response to SABAs is measured as bronchodilator drug response (BDR), which varies among racial/ethnic groups in the U.S 1, 2. However, the genetic variation that contributes to BDR is largely undefined in African Americans with asthma3\n\nObjectiveTo identify genetic variants that may contribute to differences in BDR in African Americans with asthma.\n\nMethodsWe performed a genome-wide association study of BDR in 949 African American children with asthma, genotyped with the Axiom World Array 4 (Affymetrix, Santa Clara, CA) followed by imputation using 1000 Genomes phase 3 genotypes. We used linear regression models adjusting for age, sex, body mass index and genetic ancestry to test for an association between BDR and genotype at single nucleotide polymorphisms (SNPs). To increase power and distinguish between shared vs. population-specific associations with BDR in children with asthma, we performed a meta-analysis across 949 African Americans and 1,830 Latinos (Total=2,779). Lastly, we performed genome-wide admixture mapping to identify regions whereby local African or European ancestry is associated with BDR in African Americans. Two additional populations of 416 Latinos and 1,325 African Americans were used to replicate significant associations.\n\nResultsWe identified a population-specific association with an intergenic SNP on chromosome 9q21 that was significantly associated with BDR (rs73650726, p=7.69 x 10-9). A trans-ethnic meta-analysis across African Americans and Latinos identified three additional SNPs within the intron of PRKG1 that were significantly associated with BDR (rs7903366, rs7070958, and rs7081864, p[&le;]5 x 10-8).\n\nConclusionsOur findings indicate that both population specific and shared genetic variation contributes to differences in BDR in minority children with asthma, and that the genetic underpinnings of BDR may differ between racial/ethnic groups.\n\nKey messagesO_LIA GWAS for BDR in African American children with asthma identified an intergenic population specific variant at 9q21 to be associated with increased bronchodilator drug response (BDR).\nC_LIO_LIA meta-analysis of GWAS across African Americans and Latinos identified shared genetic variants at 10q21 in the intron of PRKG1 to be associated with differences in BDR.\nC_LIO_LIFurther genetic studies need to be performed in diverse populations to identify the full set of genetic variants that contribute to BDR.\nC_LI

genomics

Ancient genomic variation underlies recent and repeated ecological adaptation

Adaptation in the wild often involves standing genetic variation (SGV), which allows rapid responses to selection on ecological timescales. However, we still know little about how the evolutionary histories and genomic distributions of SGV influence local adaptation in natural populations. Here, we address this knowledge gap using the threespine stickleback fish (Gasterosteus aculeatus) as a model. We extend the popular restriction site-associated DNA sequencing (RAD-seq) method to produce phased haplotypes approaching 700 base pairs (bp) in length at each of over 50,000 loci across the stickleback genome. Parallel adaptation in two geographically isolated freshwater pond populations consistently involved fixation of haplotypes that are identical-by-descent. In these same genomic regions, sequence divergence between marine and freshwater stickleback, as measured by dXY, reaches ten-fold higher than background levels and structures genomic variation into distinct marine and freshwater haplogroups. By combining this dataset with a de novo genome assembly of a related species, the ninespine stickleback (Pungitius pungitius), we find that this habitat-associated divergent variation averages six million years old, nearly twice the genome-wide average. The genomic variation that is involved in recent and rapid local adaptation in stickleback has actually been evolving throughout the 15-million-year history since the two species lineages split. This long history of genomic divergence has maintained large genomic regions of ancient ancestry that include multiple chromosomal inversions and extensive linked variation. These discoveries of ancient genetic variation spread broadly across the genome in stickleback demonstrate how selection on ecological timescales is a result of genome evolution over geological timescales, and vice versa.\n\nIMPACT STATEMENTAdaptation to changing environments requires a source of genetic variation. When environments change quickly, species often rely on variation that is already present - so-called standing genetic variation - because new adaptive mutations are just too rare. The threespine stickleback, a small fish species living throughout the Northern Hemisphere, is well-known for its ability to rapidly adapt to new environments. Populations living in coastal oceans are heavily armored with bony plates and spines that protect them from predators. These marine populations have repeatedly invaded and adapted to freshwater environments, losing much of their armor and changing in shape, size, color, and behavior.\n\nAdaptation to freshwater environments can occur in mere decades and probably involves lots of standing genetic variation. Indeed, one of the clearest examples we have of adaptation from standing genetic variation comes from a gene, eda, that controls the shifts in armor plating. This discovery involved two surprises that continue to shape our understanding of the genetics of adaptation. First, freshwater stickleback from across the Northern Hemisphere share the same version, or allele, of this gene. Second, the marine and freshwater alleles arose millions of years ago, even though the freshwater populations studied arose much more recently. While it has been hypothesized that other genes in the stickleback genome may share these patterns, large-scale surveys of genomic variation have been unable to test this prediction directly.\n\nHere, we use new sequencing technologies to survey DNA sequence variation across the stickleback genome for patterns like those at the eda gene. We find that every region of the genome associated with marine-freshwater genetic differences shares this pattern to some degree. Moreover, many of these regions are as old or older than eda, stretching back over 10 million years in the past and perhaps even predating the species we now call the threespine stickleback. We conclude that natural selection has maintained this variation over geological timescales and that the same alleles we observe in freshwater stickleback today are the same as those under selection in ancient, now-extinct freshwater habitats. Our findings highlight the need to understand evolution on macroevolutionary timescales to understand and predict adaptation happening in the present day.

evolutionary biology

A polygenic p factor for major psychiatric disorders

It has recently been proposed that a single dimension, called the p factor, can capture a persons liability to mental disorder. Relevant to the p hypothesis, recent genetic research has found surprisingly high genetic correlations between pairs of psychiatric disorders. Here, for the first time we compare genetic correlations from different methods and examine their support for a genetic p factor. We tested the hypothesis of a genetic p factor by using principal component analysis on matrices of genetic correlations between major psychiatric disorders estimated by three methods - family study, Genome-wide Complex Trait Analysis, and Linkage-Disequilibrium Score Regression - and on a matrix of polygenic score correlations constructed for each individual in a UK-representative sample of 7,026 unrelated individuals. All disorders loaded on a first unrotated principal component, which accounted for 57%, 43%, 34% and 19% of the variance respectively for each method. Our results showed that all four methods provided strong support for a genetic p factor that represents the pinnacle of the hierarchical genetic architecture of psychopathology.

genomics

The Genomic Formation of South and Central Asia

The genetic formation of Central and South Asian populations has been unclear because of an absence of ancient DNA. To address this gap, we generated genome-wide data from 362 ancient individuals, including the first from eastern Iran, Turan (Uzbekistan, Turkmenistan, and Tajikistan), Bronze Age Kazakhstan, and South Asia. Our data reveal a complex set of genetic sources that ultimately combined to form the ancestry of South Asians today. We document a southward spread of genetic ancestry from the Eurasian Steppe, correlating with the archaeologically known expansion of pastoralist sites from the Steppe to Turan in the Middle Bronze Age (2300-1500 BCE). These Steppe communities mixed genetically with peoples of the Bactria Margiana Archaeological Complex (BMAC) whom they encountered in Turan (primarily descendants of earlier agriculturalists of Iran), but there is no evidence that the main BMAC population contributed genetically to later South Asians. Instead, Steppe communities integrated farther south throughout the 2nd millennium BCE, and we show that they mixed with a more southern population that we document at multiple sites as outlier individuals exhibiting a distinctive mixture of ancestry related to Iranian agriculturalists and South Asian hunter-gathers. We call this group Indus Periphery because they were found at sites in cultural contact with the Indus Valley Civilization (IVC) and along its northern fringe, and also because they were genetically similar to post-IVC groups in the Swat Valley of Pakistan. By co-analyzing ancient DNA and genomic data from diverse present-day South Asians, we show that Indus Periphery-related people are the single most important source of ancestry in South Asia--consistent with the idea that the Indus Periphery individuals are providing us with the first direct look at the ancestry of peoples of the IVC--and we develop a model for the formation of present-day South Asians in terms of the temporally and geographically proximate sources of Indus Periphery-related, Steppe, and local South Asian hunter-gatherer-related ancestry. Our results show how ancestry from the Steppe genetically linked Europe and South Asia in the Bronze Age, and identifies the populations that almost certainly were responsible for spreading Indo-European languages across much of Eurasia.\n\nOne Sentence SummaryGenome wide ancient DNA from 357 individuals from Central and South Asia sheds new light on the spread of Indo-European languages and parallels between the genetic history of two sub-continents, Europe and South Asia.

genomics

A major locus for ivermectin resistance in a parasitic nematode

BackgroundInfections with helminths cause an enormous disease burden in billions of animals and plants worldwide. Large scale use of anthelmintics has driven the evolution of resistance in a number of species that infect livestock and companion animals, and there are growing concerns regarding the reduced efficacy in some human-infective helminths. Understanding the mechanisms by which resistance evolves is the focus of increasing interest; robust genetic analysis of helminths is challenging, and although many candidate genes have been proposed, the genetic basis of resistance remains poorly resolved. ResultsHere, we present a genome-wide analysis of two genetic crosses between ivermectin resistant and sensitive isolates of the parasitic nematode Haemonchus contortus, an economically important gastrointestinal parasite of small ruminants and a model for anthelmintic research. Whole genome sequencing of parental populations, and key stages throughout the crosses, identified extensive genomic diversity that differentiates populations, but after backcrossing and selection, a single genomic quantitative trait locus (QTL) localised on chromosome V was revealed to be associated with ivermectin resistance. This QTL was common between the two geographically and genetically divergent resistant populations and did not include any leading candidate genes, suggestive of a previously uncharacterised mechanism and/or driver of resistance. Despite limited resolution due to low recombination in this region, population genetic analyses and novel evolutionary models supported strong selection at this Q.TL, driven by at least partial dominance of the resistant allele, and that large resistance-associated haplotype blocks were enriched in response to selection. ConclusionsWe have described the genetic architecture and mode of ivermectin selection, revealing a major genomic locus associated with ivermectin resistance, the most conclusive evidence to date in any parasitic nematode. This study highlights a novel genome-wide approach to the analysis of a genetic cross in non-model organisms with extreme genetic diversity, and the importance of a high quality reference genome in interpreting the signals of selection so identified.

genomics