Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Behavioral individuality reveals genetic control of phenotypic variability

Variability is ubiquitous in nature and a fundamental feature of complex systems. Few studies, however, have investigated variance itself as a trait under genetic control. By focusing primarily on trait means and ignoring the effect of alternative alleles on trait variability, we may be missing an important axis of genetic variation contributing to phenotypic differences among individuals1,2. To study genetic effects on individual-to-individual phenotypic variability (or intragenotypic variability), we used a panel of Drosophila inbred lines3 and focused on locomotor handedness4, in an assay optimized to measure variability. We discovered that some lines had consistently high levels of intragenotypic variability among individuals while others had low levels. We demonstrate that the degree of variability is itself heritable. Using a genome-wide association study (GWAS) for the degree of intragenotypic variability as the phenotype across lines, we identified several genes expressed in the brain that affect variability in handedness without affecting the mean. One of these genes, Ten-a, implicated a neuropil in the central complex5 of the fly brain as influencing the magnitude of behavioral variability, a brain region involved in sensory integration and locomotor coordination6. We have validated these results using genetic deficiencies, null alleles, and inducible RNAi transgenes. This study reveals the constellation of phenotypes that can arise from a single genotype and it shows that different genetic backgrounds differ dramatically in their propensity for phenotypic variability. Because traditional mean-focused GWASs ignore the contribution of variability to overall phenotypic variation, current methods may miss important links between genotype and phenotype.\n\nAbbreviationsDGRP: Drosophila Genome Reference Panel; ANOMV: Analysis of means for variance; QTL: quantitative trait loci, GWAS: genome wide association study, MAD: median absolute deviation. GWAS: genome wide association study, CI: confidence interval.

Genetics

The genetic architecture of neurodevelopmental disorders

Neurodevelopmental disorders include rare conditions caused by identified single mutations, such as Fragile X, Down and Angelman syndromes, and much more common clinical categories such as autism, epilepsy and schizophrenia. These common conditions are all highly heritable but their genetics is considered to be \"complex\". In fact, this sharp dichotomy in genetic architecture between rare and common disorders may be largely artificial. On the one hand, much of the apparent complexity in the genetics of common disorders may derive from underlying genetic heterogeneity, which has remained obscure until recently. On the other hand, even for supposedly Mendelian conditions, the relationship between single mutations and clinical phenotypes is rarely simple. The categories of monogenic and complex disorders may therefore merge across a continuum, with some mutations being strongly associated with specific syndromes and others having a more variable outcome, modified by the presence of additional genetic variants.

Genetics

Genetic Analysis of Substrain Divergence in NOD Mice

ABSTRACTThe NOD mouse is a polygenic model for type 1 diabetes that is characterized by insulitis, a leukocytic infiltration of the pancreatic islets. During ~35 years since the original inbred strain was developed in Japan, NOD substrains have been established at different laboratories around the world. Although environmental differences among NOD colonies capable of impacting diabetes incidence have been recognized, differences arising from genetic divergence have not previously been analyzed. We use both Mouse Diversity Array and Whole Exome Capture Sequencing platforms to identify genetic differences distinguishing 5 NOD substrains. We describe 64 SNPs, and 2 short indels that differ in coding regions of the 5 NOD substrains. A 100 kb deletion on Chromosome 3 distinguishes NOD/ShiLtJ and NOD/ShiLtDvs from 3 other substrains, while a 111 kb deletion in the Icam2 gene on Chromosome 11 is unique to the NOD/ShiLtDvs genome. The extent of genetic divergence for NOD substrains is compared to similar studies for C57BL6 and BALB/c substrains. As mutations are fixed to homozygosity by continued inbreeding, significant differences in substrain phenotypes are to be expected. These results emphasize the importance of using embryo freezing methods to minimize genetic drift within substrains and of applying appropriate genetic nomenclature to permit substrain recognition when one is used.

Genetics

Scaling probabilistic models of genetic variation to millions of humans

One of the major goals of population genetics is to quantitatively understand variation of genetic polymorphisms among individuals. To this end, researchers have developed sophisticated statistical methods to capture the complex population structure that underlies observed genotypes in humans, and such methods have been effective for analyzing modestly sized genomic data sets. However, the number of genotyped humans has grown significantly in recent years, and it is accelerating. In aggregate about 1M individuals have been genotyped to date. Analyzing these data will bring us closer to a nearly complete picture of human genetic variation; but existing methods for population genetics analysis do not scale to data of this size. To solve this problem we developed TeraStructure. TeraStructure is a new algorithm to fit Bayesian models of genetic variation in human populations on tera-sample-sized data sets (1012 observed genotypes, e.g., 1M individuals at 1M SNPs). It is a principled approach to Bayesian inference that iterates between subsampling locations of the genome and updating an estimate of the latent population structure of the individuals. On data sets of up to 2K individuals, TeraStructure matches the existing state of the art in terms of both speed and accuracy. On simulated data sets of up to 10K individuals, TeraStructure is twice as fast as existing methods and has higher accuracy in recovering the latent population structure. On genomic data simulated at the tera-sample-size scales, TeraStructure continues to be accurate and is the only method that can complete its analysis.\n\nSoftwareTeraStructure is available for download at https://github.com/premgopalan/terastructure.\n\nFundingThis research was supported in part by NIH grant R01 HG006448 and ONR grant N00014-12-1-0764.

Genetics

Genetics of intra-species variation in avoidance behavior induced by a thermal stimulus in C. elegans

Individuals within a species vary in their responses to a wide range of stimuli, partly as a result of differences in their genetic makeup. Relatively little is known about the genetic and neuronal mechanisms contributing to diversity of behavior in natural populations. By studying animal-to-animal variation in innate avoidance behavior to thermal stimuli in the nematode Caenorhabditis elegans, we uncovered genetic principles of how different components of a behavioral response can be altered in nature to generate behavioral diversity. Using a thermal pulse assay, we uncovered heritable variation in responses to a transient temperature increase. Quantitative trait locus mapping revealed that separate components of this response were controlled by distinct genomic loci. The loci we identified contributed to variation in components of thermal pulse avoidance behavior in an additive fashion. Our results show that the escape behavior induced by thermal stimuli is composed of simpler behavioral components that are influenced by at least six distinct genetic loci. The loci that decouple components of the escape behavior reveal a genetic system that allows independent modification of behavioral parameters. Our work sets the foundation for future studies of evolution of innate behaviors at the molecular and neuronal level.

Genetics

An Atlas of Genetic Correlations across Human Diseases and Traits

Identifying genetic correlations between complex traits and diseases can provide useful etiological insights and help prioritize likely causal relationships. The major challenges preventing estimation of genetic correlation from genome-wide association study (GWAS) data with current methods are the lack of availability of individual genotype data and widespread sample overlap among meta-analyses. We circumvent these difficulties by introducing a technique for estimating genetic correlation that requires only GWAS summary statistics and is not biased by sample overlap. We use our method to estimate 300 genetic correlations among 25 traits, totaling more than 1.5 million unique phenotype measurements. Our results include genetic correlations between anorexia nervosa and schizophrenia, anorexia and obesity and associations between educational attainment and several diseases. These results highlight the power of genome-wide analyses, since there currently are no genome-wide significant SNPs for anorexia nervosa and only three for educational attainment.

Genomics

The genetics of Bene Israel from India reveals both substantial Jewish and Indian ancestry

The Bene Israel Jewish community from West India is a unique population whose history before the 18th century remains largely unknown. Bene Israel members consider themselves as descendants of Jews, yet the identity of Jewish ancestors and their arrival time to India are unknown, with speculations on arrival time varying between the 8th century BCE and the 6th century CE. Here, we characterize the genetic history of Bene Israel by collecting and genotyping 18 Bene Israel individuals. Combining with 486 individuals from 41 other Jewish, Indian and Pakistani populations, and additional individuals from worldwide populations, we conducted comprehensive genome-wide analyses based on FST, principal component analysis, ADMIXTURE, identity-by-descent sharing, admixture linkage disequilibrium decay, haplotype sharing and allele sharing autocorrelation decay, as well as contrasted patterns between the X chromosome and the autosomes. The genetics of Bene Israel individuals resemble local Indian populations, while at the same time constituting a clearly separated and unique population in India. They are unique among Indian and Pakistani populations we analyzed in sharing considerable genetic ancestry with other Jewish populations. Putting together the results from all analyses point to Bene Israel being an admixed population with both Jewish and Indian ancestry, with the genetic contribution of each of these ancestral populations being substantial. The admixture took place in the last millennium, about 19-33 generations ago. It involved Middle-Eastern Jews and was sex-biased, with more male Jewish and local female contribution. It was followed by a population bottleneck and high endogamy, which can lead to increased prevalence of recessive diseases in this population. This study provides an example of how genetic analysis advances our knowledge of human history in cases where other disciplines lack the relevant data to do so.

Genetics

Differential methylation between ethnic sub-groups reflects the effect of genetic ancestry and environmental exposures

In clinical practice and biomedical research populations are often divided categorically into distinct racial/ethnic groups. In reality, these categories, which are based on social rather than biological constructs, comprise diverse groups with highly heterogeneous histories, cultures, traditions, religions, social and environmental exposures and ancestral backgrounds. Their use is thus widely debated and genetic ancestry has been suggested as a complement or alternative to this categorization. However, few studies have examined the relative contributions of racial/ethnic identity, genetic ancestry, and environmental exposures on well-established and fundamental biological processes. We examined the associations between ethnicity, ancestry, and environmental exposures and DNA methylation. We typed over 450,000 CpG sites in primary whole blood of 573 individuals of diverse Hispanic descent who also had high-density genotype data. We found that both self-identified ethnicity and genetically determined ancestry were significantly associated with methylation levels at a large number of CpG sites (916 and 194, respectively). Among loci differentially methylated between ethnic groups, a median of 75.7% (IQR 45.8% to 92%) of the variance in methylation associated with ethnicity could be accounted for by shared genomic ancestry accounts. We also found significant enrichment (p = 4.2 x 10-64) of ethnicity-associated sites amongst loci previously associated with environmental and social exposures, particularly maternal smoking during pregnancy. Our study suggests that although differential methylation between ethnic groups can be partially explained by the shared genetic ancestry, a significant effect of ethnicity is likely due to environmental, social, or cultural factors, which differ between ethnic groups.\n\nOne Sentence SummaryIn order to better understand the role of ethnic self-identification and genetically determined ancestry in biomedical outcomes, we explore their relative contributions to variation in methylation, a fundamental biological process.\n\nSources of FundingThis research was supported in part by the Sandler Family Foundation, the American Asthma Foundation, National Institutes of Health (P60 MD006902, R01 HL117004, R21ES24844, U54MD009523, R01 ES015794, R01 HL088133, M01 RR000083, R01 HL078885, R01 HL104608, U19 AI077439, M01 RR00188, U01 HG009080, and R01 HL135156), ARRA grant RC2 HL101651, and TRDRP 24RT-0025; EGB was supported in part through grants from the Flight Attendant Medical Research Institute (FAMRI), and NIH (K23 HL004464); NZ was supported in part by an NIH career development award from the NHLBI (K25HL121295). JMG was supported in part by NIH Training Grant T32 (T32GM007546) and career development awards from the NHLBI (K23HL111636) and NCATS (KL2TR000143) as well as the Hewett Fellowship; N.T. was supported in part by an institutional training grant from the NIGMS (T32-GM007546) and career development awards from the NHLBI (K12-HL119997 and K23-HL125551), Parker B. Francis Fellowship Program, and the American Thoracic Society; CRG was supported in part by NIH Training Grant T32 (GM007175) and the UCSF Chancellors Research Fellowship and Dissertation Year Fellowship; RK was supported with a career development award from the NHLBI (K23HL093023); HJF was supported in part by the GCRC (RR00188); PCA was supported in part by the Ernest S. Bazley Grant; MAS was supported in part by 1R01HL128439-01. This publication was supported by various institutes within the National Institutes of Health. Its contents are solely the responsibility of the authors and do not necessarily represent the official views of the NIH.

Genetics

Genetic Influences on Hormonal Markers of Chronic HPA Function in Human Hair

Cortisol is the primary output of the hypothalamic-pituitary-adrenal (HPA) axis and is central to the human biological stress response, with wide-ranging effects on physiological function and psychiatric health. In both humans and animals, cortisol is frequently studied as a biomarker for exposure to environmental stress. Relatively little attention has been paid to the possible role of genetic variation in heterogeneity in chronic cortisol, in spite of well-studied biological pathways of glucocorticoid function. Using recently developed technology, hair samples can now be used to measure accumulation of cortisol over several months. In contrast to more conventional salivary measures, hair cortisol is not influenced by diurnal variation or transient hormonal reactivity. In an ethnically and socioeconomically diverse sample of 1 070 child and adolescent twins and multiples from 556 unique families, we estimated genetic and environmental influences on hair concentrations of cortisol and its inactive metabolite, cortisone. We identified sizable genetic influences on cortisol that decrease with age, concomitant with genetic influences on cortisone that increase with age. Shared environmental influences on cortisol and cortisone were modest and, for cortisol, decreased with age. Twin-specific, non-shared environmental contributions to cortisol and cortisone became increasingly correlated with age. We find some evidence for sex differences in the biometric contributions to cortisol, but no strong evidence for main or moderating effects of family socioeconomic status on cortisol or cortisone. This study constitutes the first genetic study of hormone concentrations in human hair, and provides the most definitive characterization to-date of age and socioeconomic influences on hair cortisol.

Genetics

Maternal Genetic Ancestry and Legacy of 10th Century AD Hungarians

The ancient Hungarians originated from the Ural region in todays central Russia and migrated across the Eastern European steppe, according to historical sources. The Hungarians conquered the Carpathian Basin 895-907 AD, and admixed with the indigenous communities.\n\nHere we present mitochondrial DNA results from three datasets: one from the Avar period (7th -9th centuries) of the Carpathian Basin (n = 31); one from the Hungarian conquest-period (n=76); and a completion of the published 10th-12th century Hungarian-Slavic contact zone dataset by four samples. We compare these mitochondrial DNA hypervariable segment sequences and haplogroup results with published ancient and modern Eurasian data. Whereas the analyzed Avars represents a certain group of the Avar society that shows East and South European genetic characteristics, the Hungarian conquerors maternal gene pool is a mixture of West Eurasian and Central and North Eurasian elements. Comprehensively analyzing the results, both the linguistically recorded Finno-Ugric roots and historically documented Turkic and Central Asian influxes had possible genetic imprints in the conquerors genetic composition. Our data allows a complex series of historic and population genetic events before the formation of the medieval population of the Carpathian Basin, and the maternal genetic continuity between 10th- 12th century and modern Hungarians.

Genetics

Do Regional Brain Volumes and Major Depressive Disorder Share Genetic Architecture: a study in Generation Scotland (n=19,762), UK Biobank (n=24,048) and the English Longitudinal Study of Ageing (n=5,766)

Major depressive disorder (MDD) is a heritable and highly debilitating condition. It is commonly associated with subcortical volumetric abnormalities, the most replicated of these being reduced hippocampal volume. Using the most recent published data from ENIGMA consortiums genome-wide association study (GWAS) of regional brain volume, we sought to test whether there is shared genetic architecture between 8 subcortical brain volumes and MDD. Using LD score regression utilising summary statistics from ENIGMA and the Psychiatric Genomics Consortium, we demonstrated that hippocampal volume was positively genetically correlated with MDD (rG=0.46, P=0.02), although this did not survive multiple comparison testing. None of other six brain regions studied were genetically correlated and amygdala volume heritability was too low for analysis. We also generated polygenic risk scores (PRS) to assess potential pleiotropy between regional brain volumes and MDD in three cohorts (Generation Scotland; Scottish Family Health Study (n=19,762), UK Biobank (n=24,048) and the English Longitudinal Study of Ageing (n=5,766). We used logistic regression to examine volumetric PRS and MDD and performed a meta-analysis across the three cohorts. No regional volumetric PRS demonstrated significant association with MDD or recurrent MDD. In this study we provide some evidence that hippocampal volume and MDD may share genetic architecture, albeit this did not survive multiple testing correction and was in the opposite direction to most reported phenotypic correlations. We therefore found no evidence to support a shared genetic architecture for MDD and regional subcortical volumes.

Genetics

Accounting for genetic interactions is necessary for accurate prediction of extreme phenotypic values of quantitative traits in yeast

Experiments in model organisms report abundant genetic interactions underlying biologically important traits, whereas quantitative genetics theory predicts, and data support, that most genetic variance in populations is additive. Here we describe networks of capacitating genetic interactions that contribute to quantitative trait variation in a large yeast intercross population. The additive variance explained by individual loci in a network is highly dependent on the allele frequencies of the interacting loci. Modeling of phenotypes for multi-locus genotype classes in the epistatic networks is often improved by accounting for the interactions. We discuss the implications of these results for attempts to dissect genetic architectures and to predict individual phenotypes and long-term responses to selection.

Genetics

Genetic correlations with climate variables suggest Caenorhabditis elegans natural niche preferences

Species inhabit a variety of environmental niches, and the adaptation to a particular niche is often controlled by genetic factors, including gene-by-environment interactions. The genes that vary in order to regulate the ability to colonize a niche are often difficult to identify, especially in the context of complex ecological systems and in experimentally uncontrolled natural environments. Quantitative genetic approaches provide an opportunity to investigate correlations between genetic factors and environmental parameters that might define a niche. Previously, we have shown how a collection of 208 whole-genome sequenced wild Caenorhabditis elegans can facilitate association mapping approaches. To correlate climate parameters with the variation found in this collection of wild strains, we used geographic data to exhaustively curate daily weather measurements in short-term (three month), middle-term (one year), and long-term (three year) durations surrounding the data of strain isolation. These climate parameters were then used as quantitative traits in the mapping approaches. We identified 10 QTL underlying variation in three traits: elevation, relative humidity, and average temperature. We then performed statistical analyses to further narrow the genomic interval of interest to identify gene candidates with variants potentially underlying phenotypic differences. Additionally, we performed two-strain competition assays at high and low temperatures to validate a QTL for temperature preference and found suggestive evidence that genotypes might be adapted to particular temperatures.\n\n100-word summary for G3Quantitative genetic approaches provide an opportunity to investigate correlations between genetic factors and environmental parameters that might define a niche, but these genes are difficult to identify, especially in the context of complex ecological systems. Here, we used a collection of 152 sequenced wild Caenorhabditis elegans to correlate climate parameters with the variation found in this collection of wild strains. We identified 10 QTL in five traits, including elevation, relative humidity, and temperature. Additionally, we performed competition assays to validate a QTL for temperature preference and found suggestive evidence that genotypes might be adapted to particular temperatures.

Genetics

SweGen: A whole-genome map of genetic variability in a cross-section of the Swedish population

Here we describe the SweGen dataset, a high-quality map of genetic variation in the Swedish population. This data represents a basic resource for clinical genetics laboratories as well as for sequencing-based association studies, by providing information on the frequencies of genetic variants in a cohort that is well matched to national patient cohorts. To select samples for this study, we first examined the genetic structure of the Swedish population using high-density SNP-array data from a nation-wide population based cohort of over 10,000 individuals. From this sample collection, 1,000 individuals, reflecting a cross-section of the population and capturing the main genetic structure, were selected for whole genome sequencing (WGS). Analysis pipelines were developed for automated alignment, variant calling and quality control of the sequencing data. This resulted in a whole-genome map of aggregated variant frequencies in the Swedish population that we hereby release to the scientific community.

genetics

SOFIA: an R package for enhancing genetic visualization with Circos

Visualization of data from any stage of genetic and genomic research is one of the most useful approaches for detecting potential errors, ensuring accuracy and reproducibility, and presentation of the resulting data. Currently software such as Circos, ClicO FS, and RCircos, among others, provide tools for plotting a variety of genetic data types in a concise manner for data exploration and presentation. However, each of the programs have one or more disadvantages that limit their usability in data exploration or construction of publication quality figures, such as inflexibility in formatting and configuration, reduced image quality, lack of potential for automation, or requirements of high-level computational expertise. Therefore, we developed the R package SOFIA, which leverages the capabilities of Circos by manipulating data, preparing configuration files, and running the Perl-native Circos directly from the R environment with minimal user intervention. The advantages of integrating both R and Circos into SOFIA are numerous. R is a very powerful, mid-level programming language widely used among the genetic and genomic research community, while Circos has proven to be a novel software for arranging genomic data to create aesthetical publication quality circular figures. Producing Circos figures in R with SOFIA is simple, requires minimal coding experience, even for complex figures that incorporate high-dimensional genetic information, and allows simultaneous analysis and visual exploration of genomic and genetic data in a single programming environment.

genetics

A genetic risk score to guide age-specific, personalized prostate cancer screening

BackgroundProstate-specific-antigen (PSA) screening resulted in reduced prostate cancer (PCa) mortality in a large clinical trial, but due to a high false-positive rate, among other concerns, many guidelines do not endorse universal screening and instead recommend an individualized decision based on each patients risk. Genetic risk may provide key information to guide the decisions of whether and at what age to screen an individual man for PCa.\n\nMethodsGenotype, PCa status, and age from 34,444 men of European ancestry from the PRACTICAL consortium database were analyzed to select single-nucleotide polymorphisms (SNPs) associated with prostate cancer diagnosis. These SNPs were then incorporated into a survival analysis to estimate their effects on age at PCa diagnosis. The resulting polygenic hazard score (PHS) is an assessment of individual genetic risk. The final model was validated in an independent dataset comprised of 6,417 men with screening PSA and genotype data. PHS was calculated for these men to test for prediction of PCa-free survival. PHS was also combined with age-specific PCa incidence data from the U.S. population to generate a PCa-Risk (PCaR) age that relates a given mans risk to that of the population average. PHS and PCaR age were evaluated for prediction of positive predictive value (PPV) of PSA screening.\n\nFindingsPHS calculated from 54 SNPs was very highly predictive of age at PCa diagnosis for men in the validation set (p =10-53). PPV of PSA screening varied from 0.18 to 0.52 for men with low and high genetic risk, respectively. PHS modulates PCa-free survival curves by an estimated 20 years between men in the 1st or 99th percentiles of genetic risk.\n\nInterpretationPolygenic hazard scores give personalized genetic risk estimates and can inform the decisions of whether and at what age to screen a man for PCa.\n\nFundingDepartment of Defense #W81XWH-13-1-0391

genetics

Genetic and transgenic reagents for Drosophila simulans, D.mauritiana, D. yakuba, D. santomea and D. virilis

Species of the Drosophila melanogaster species subgroup, including the species D. simulans, D. mauritiana, D. yakuba, and D. santomea, have long served as model systems for studying evolution. Studies in these species have been limited, however, by a paucity of genetic and transgenic reagents. Here we describe a collection of transgenic and genetic strains generated to facilitate genetic studies within and between these species. We have generated many strains of each species containing mapped piggyBac transposons including an enhanced yellow fluorescent protein gene expressed in the eyes and a phiC31 attP site-specific integration site. We have tested a subset of these lines for integration efficiency and reporter gene expression levels. We have also generated a smaller collection of other lines expressing other genetically encoded fluorescent molecules in the eyes and a number of other transgenic reagents that will be useful for functional studies in these species. In addition, we have mapped the insertion locations of 58 transposable elements in D. virilis that will be useful for genetic mapping studies.

genetics

An efficient Bayesian meta-analysis approach for studying cross-phenotype genetic associations

Simultaneous analysis of genetic associations with multiple phenotypes may reveal shared genetic susceptibility across traits (pleiotropy). For a locus exhibiting overall pleiotropy, it is important to identify which specific traits underlie this association. We propose a Bayesian meta-analysis approach (termed CPBayes) that uses summary-level data across multiple phenotypes to simultaneously measure the evidence of aggregate-level pleiotropic association and estimate an optimal subset of traits associated with the risk locus. This method uses a unified Bayesian statistical framework based on a spike and slab prior. CPBayes performs a fully Bayesian analysis by employing the Markov chain Monte Carlo (MCMC) technique Gibbs sampling. It takes into account heterogeneity in the size and direction of the genetic effects across traits. It can be applied to both cohort data and separate studies of multiple traits having overlapping or non-overlapping subjects. Simulations show that CPBayes produces a substantially better accuracy in the selection of associated traits underlying a pleiotropic signal than the subset-based meta-analysis ASSET. We used CPBayes to undertake a genome-wide pleiotropic association study of 22 traits in the large Kaiser GERA cohort and detected nine independent pleiotropic loci associated with at least two phenotypes. This includes a locus at chromosomal region 1q24.2 which exhibits an association simultaneously with the risk of five different diseases: Dermatophytosis, Hemorrhoids, Iron Deficiency, Osteoporosis, and Peripheral Vascular Disease. The GERA cohort analysis suggests that CPBayes is more powerful than ASSET with respect to detecting independent pleiotropic variants. We provide an R-package CPBayes implementing the proposed method.\n\nAuthor SummaryGenome-wide association studies (GWASs) have highlighted shared genetic susceptibility to various human diseases (pleiotropy). We propose a Bayesian meta-analysis method CPBayes that simultaneously evaluates the evidence of aggregate-level pleiotropic association and selects an optimal subset of associated traits underlying a pleiotropic signal. CPBayes analyzes pleiotropy using summary-level data across a wide range of studies for two or more phenotypes - separate GWASs with or without shared subjects, cohort study for multiple traits. It performs a fully Bayesian analysis and offers various flexibilities in the inference. In addition to parameters of primary interest (e.g., the measures of overall pleiotropic association, the optimal subset of associated traits), it provides additional interesting insights into a pleiotropic signal (e.g., the trait-specific posterior probability of association, the credible interval of unknown true genetic effects). Using computer simulations and a real data application to the large Kaiser GERA cohort, we demonstrate that CPBayes offers substantially better accuracy while selecting the non-null traits compared to a well known subset-based meta analysis ASSET. In the GERA cohort analysis, CPBayes detected a larger number of independent pleiotropic variants than ASSET. We provide a user-friendly R-package CPBayes for general use.

genetics