Search bioRxivSearch

Biology subjects

Sofer, T.

Publications and source records attributed to Sofer, T..

8 recordsLinked to original sources

Genome-wide association analysis of excessive daytime sleepiness identifies 42 loci that suggest phenotypic subgroups

Excessive daytime sleepiness (EDS) affects 10-20% of the population and is associated with substantial functional deficits. We identified 42 loci for self-reported EDS in GWAS of 452,071 individuals from the UK Biobank, with enrichment for genes expressed in brain tissues and in neuronal transmission pathways. We confirmed the aggregate effect of a genetic risk score of 42 SNPs on EDS in independent Scandinavian cohorts and on other sleep disorders (restless leg syndrome, insomnia) and sleep traits (duration, chronotype, accelerometer-derived sleep efficiency and daytime naps or inactivity). Strong genetic correlations were also seen with obesity, coronary heart disease, psychiatric diseases, cognitive traits and reproductive ageing. EDS variants clustered into two predominant composite phenotypes - sleep propensity and sleep fragmentation - with the former showing stronger evidence for enriched expression in central nervous system tissues, suggesting two unique mechanistic pathways. Mendelian randomization analysis indicated that higher BMI is causally associated with EDS risk, but EDS does not appear to causally influence BMI.

genomics

Epigenome-wide association analysis of daytime sleepiness in the Multi-Ethnic Study of Atherosclerosis reveals African-American specific associations

Study ObjectivesExcessive daytime sleepiness (EDS) is a consequence of inadequate sleep, or of a primary disorder of sleep-wake control. Population variability in prevalence of EDS and susceptibility to EDS are likely due to genetic and biological factors as well as social and environmental influences. Epigenetic modifications (such as DNA methylation-DNAm) are potential influences on a range of health outcomes. Here, we explored the association between DNAm and daytime sleepiness quantified by the Epworth Sleepiness Scale (ESS).\n\nMethodsWe performed multi-ethnic and ethnic-specific epigenome-wide association studies for DNAm and ESS in 619 individuals from the Multi-Ethnic Study of Atherosclerosis. Replication was assessed in the Cardiovascular Health Study (CHS). Genetic variants in genes proximal to ESS-associated DNAm were analyzed to identify methylation quantitative trait loci and followed with replication of genotype-sleepiness associations in the UK Biobank.\n\nResults61 methylation sites were associated with ESS (FDR [≤] 0.1) in African Americans only, including an association in KCTD5, a gene strongly implicated in sleep. One association (cg26130090) replicated in CHS African Americans (p-value 0.0004). We identified a sleepiness-associated methylation site in the gene RAI1, a gene associated with sleep and circadian phenotypes. In a follow-up analysis, a genetic variant within RAI1 associated with both DNAm and sleepiness score. The variants association with sleepiness was replicated in the UK Biobank.\n\nConclusionsOur analysis identified methylation sites in multiple genes that may be implicated in EDS. These sleepiness-methylation associations were specific to African Americans. Future work is needed to identify mechanisms driving ancestry-specific methylation effects.\n\nStatement of SignificanceExcessive daytime sleepiness is associated with negative health outcomes such as reduction in quality of life, increased workplace accidents, and cardiovascular mortality. There are race/ethnic disparities in excessive daytime sleepiness, however, the environmental and biological mechanisms for these differences are not yet understood. We performed an association analysis of DNA methylation, measured in monocytes, and daytime sleepiness within a racially diverse study population. We detected numerous DNA methylation markers associated with daytime sleepiness in African Americans, but not in European and Hispanic Americans. Future work is required to elucidate the pathways between DNA methylation, sleepiness, and related behavioral/environmental exposures.

genomics

Efficient variant set mixed model association tests for continuous and binary traits in large-scale whole genome sequencing studies

With advances in Whole Genome Sequencing (WGS) technology, more advanced statistical methods for testing genetic association with rare variants are being developed. Methods in which variants are grouped for analysis are also known as variant-set, gene-based, and aggregate unit tests. The burden test and Sequence Kernel Association Test (SKAT) are two widely used variant-set tests, which were originally developed for samples of unrelated individuals and later have been extended to family data with known pedigree structures. However, computationally-efficient and powerful variant-set tests are needed to make analyses tractable in large-scale WGS studies with complex study samples. In this paper, we propose the variant-Set Mixed Model Association Tests (SMMAT) for continuous and binary traits using the generalized linear mixed model framework. These tests can be applied to large-scale WGS studies involving samples with population structure and relatedness, such as in the National Heart, Lung, and Blood Institutes Trans-Omics for Precision Medicine (TOPMed) program. SMMAT tests share the same null model for different variant sets, and a virtue of this null model, which includes covariates only, is that it needs to be only fit once for all tests in each genome-wide analysis. Simulation studies show that all the proposed SMMAT tests correctly control type I error rates for both continuous and binary traits in the presence of population structure and relatedness. We also illustrate our tests in a real data example of analysis of plasma fibrinogen levels in the TOPMed program (n = 23,763), using the Analysis Commons, a cloud-based computing platform.

genetics

GWAS of QRS Duration Identifies New Loci Specific to Hispanic/Latino Populations

BackgroundThe electrocardiographically quantified QRS duration measures ventricular depolarization and conduction. QRS prolongation has been associated with poor heart failure prognosis and cardiovascular mortality, including sudden death. While previous genome-wide association studies (GWAS) have identified 32 QRS SNPs across 26 loci among European, African, and Asian-descent populations, the genetics of QRS among Hispanics/Latinos has not been previously explored.\n\nMethodsWe performed a GWAS of QRS duration among Hispanic/Latino ancestry populations (n=15,124) from four studies using 1000 Genomes imputed genotype data (adjusted for age, sex, global ancestry, clinical and study-specific covariates). Study-specific results were combined using fixed-effects, inverse variance-weighted meta-analysis.\n\nResultsWe identified six loci associated with QRS (P<5x10-8), including two novel loci: MYOCD, a nuclear protein expressed in the heart, and SYT1, an integral membrane protein. The top association in the MYOCD locus, intronic SNP rs16946539, was found in Hispanics/Latinos with a minor allele frequency (MAF) of 0.04, but is monomorphic in European and African descent populations. The most significant QRS duration association was for intronic SNP rs3922344 (P= 8.56x10-26) in SCN5A/SCN10A. Three additional previously identified loci, CDKN1A, VTI1A, and HAND1, also exceeded the GWAS significance threshold among Hispanics/Latinos. A total of 27 of 32 previously identified QRS duration SNPs were shown to generalize in Hispanics/Latinos.\n\nConclusionsOur QRS duration GWAS, the first in Hispanic/Latino populations, identified two new loci, underscoring the utility of extending large scale genomic studies to currently under-examined populations.

genetics

A Fully-Adjusted Two-Stage Procedure for Rank Normalization in Genetic Association Studies

When testing genotype-phenotype associations using linear regression, departure of the trait distribution from normality can impact both Type I error rate control and statistical power, with worse consequences for rarer variants. While it has been shown that applying a rank-normalization transformation to trait values before testing may improve these statistical properties, the factor driving them is not the trait distribution itself, but its residual distribution after regression on both covariates and genotype. Because genotype is expected to have a small effect (if any) investigators now routinely use a two-stage method, in which they first regress the trait on covariates, obtain residuals, rank-normalize them, and then secondly use the rank-normalized residuals in association analysis with the genotypes. Potential confounding signals are assumed to be removed at the first stage, so in practice no further adjustment is done in the second stage. Here, we show that this widely-used approach can lead to tests with undesirable statistical properties, due to both a combination of a mis-specified mean-variance relationship, and remaining covariate associations between the rank-normalized residuals and genotypes. We demonstrate these properties theoretically, and also in applications to genome-wide and whole-genome sequencing association studies. We further propose and evaluate an alternative fully-adjusted two-stage approach that adjusts for covariates both when residuals are obtained, and in the subsequent association test. This method can reduce excess Type I errors and improve statistical power.

genetics

Generalizing Genetic Risk Scores from Europeans to Hispanics/Latinos

Genetic risk scores (GRSs) are weighted sums of risk allele counts of single nucleotide polymorphisms (SNPs) associated with a disease or trait. Construction of GRSs is typically based on published results from Genome-Wide Association Studies (GWASs), the majority of which have been performed in large populations of European ancestry (EA) individuals. While many genotype-trait associations have been shown to generalize from EA populations to other populations, such as Hispanics/Latinos, the optimal choice of SNPs and weights for GRSs may differ between populations due to different linkage disequilibrium (LD) and allele frequency patterns. This is further complicated by the fact that different Hispanic/Latino populations may have different admixture patterns, so that LD and allele frequency patterns may not be the same among non-EA populations. Here, we compare various approaches for GRS construction, using GWAS results from both large EA studies and a smaller study in Hispanics/Latinos, the Hispanic Community Health Study/Study of Latinos (HCHS/SOL, n = 12, 803). We consider multiple ways to select SNPs from association regions and to calculate the SNP weights. We study the performance of the resulting GRSs in an independent study of Hispanics/Latinos from the Woman Health Initiative (WHI, n = 3, 582). We support our investigation with simulation studies of potential genetic architectures in a single locus. We observed that selecting variants based on EA GWASs generally performs well, as long as SNP weights are calculated using Hispanics/Latinos GWASs, or using the meta-analysis of EA and Hispanics/Latinos GWASs. The optimal approach depends on the genetic architecture of the trait.

genetics

Multiethnic Meta-analysis Identifies New Loci for Pulmonary Function

Nearly 100 loci have been identified for pulmonary function, almost exclusively in studies of European ancestry populations. We extend previous research by meta-analyzing genome-wide association studies of 1000 Genomes imputed variants in relation to pulmonary function in a multiethnic population of 90,715 individuals of European (N=60,552), African (N=8,429), Asian (N=9,959), and Hispanic/Latino (N=11,775) ethnicities. We identified over 50 novel loci at genome-wide significance in ancestry-specific and/or multiethnic meta-analyses. Recent fine mapping methods incorporating functional annotation, gene expression, and/or differences in linkage disequilibrium between ethnicities identified potential causal variants and genes at known and newly identified loci. Sixteen of the novel genes encode proteins with predicted or established drug targets, including KCNK2 and CDK12.

genetics

Integrated Computing And Tracking System For Centralized High-Throughput Genetic Analysis: A Case Study

The Genetic Analysis Center (GAC) of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) developed an Integrated Computing and Tracking system (ICT) in order to perform genome-wide and other genetic association studies automatically and efficiently, while documenting all analysis specifications. This system provides easy-to-use analysis set-up and computing procedures, automatic reports, and analysis search functionality due to integration with an on-site database. In this paper we describe the ICT and demonstrate how it satisfies key principles of reproducible research, while respecting constraints and challenges arising from using very large, restricted access, human-subjects data. This case study may benefit other groups that have similar requirements for high-throughput analysis execution and management.

bioinformatics