Search bioRxivSearch

Biology subjects

Albert Hofman

Publications and source records attributed to Albert Hofman.

7 recordsLinked to original sources

Partial derivatives meta-analysis: pooled analyses when individual participant data cannot be shared

Joint analysis of data from multiple studies in collaborative efforts strengthens scientific evidence, with the gold standard approach being the pooling of individual participant data (IPD). However, sharing IPD often has legal, ethical, and logistic constraints for sensitive or high-dimensional data, such as in clinical trials, observational studies, and large-scale omics studies. Therefore, meta-analysis of study-level effect estimates is routinely done, but this compromises on statistical power, accuracy, and flexibility. Here we propose a novel meta-analytical approach, named partial derivatives meta-analysis, that is mathematically equivalent to using IPD, yet only requires the sharing of aggregate data. It not only yields identical results as pooled IPD analyses, but also allows post-hoc adjustments for covariates and stratification without the need for site-specific re-analysis. Thus, in case that IPD cannot be shared, partial derivatives meta-analysis still produces gold standard results, which can be used to better inform guidelines and policies on clinical practice.

Bioinformatics

HASE:Framework for efficient high-dimensional association analyses

Large-scale data collection and processing have facilitated scientific discoveries in fields such as genomics and imaging, but cross-investigations between multiple big datasets remain impractical. Computational requirements of high-dimensional association studies are often too demanding for individual sites. Additionally, the sheer size of intermediate results is unfit for collaborative settings where summary statistics are exchanged for meta-analyses. Here we introduce the HASE framework to perform high-dimensional association studies with dramatic reduction in both computational burden and storage requirements of intermediate results. We implemented a novel meta-analytical method that yields identical power as pooled analyses without the need of sharing individual participant data. The efficiency of the framework is illustrated by associating 9 million genetic variants with 1.5 million brain imaging voxels in three cohorts (total N=4,034) followed by meta-analysis, on a standard computational infrastructure. These experiments indicate that HASE facilitates high-dimensional association studies enabling large multicenter association studies for future discoveries.

Genetics

Maternal genome-wide association study identifies a fasting glucose variant associated with offspring birth weight

Genome-wide association studies (GWAS) of birth weight have focused on fetal genetics, while relatively little is known about how maternal genetic variation influences fetal growth. We aimed to identify maternal genetic variants associated with birth weight that could highlight potentially relevant maternal determinants of fetal growth.\n\nWe meta-analysed GWAS data on up to 8.7 million SNPs in up to 86,577 women of European descent from the Early Growth Genetics (EGG) Consortium and the UK Biobank. We used structural equation modelling (SEM) and analyses of mother-child pairs to quantify the separate maternal and fetal genetic effects.\n\nMaternal SNPs at 10 loci (MTNR1B, HMGA2, SH2B3, KCNAB1, L3MBTL3, GCK, EBF1, TCF7L2, ACTL9 and CYP3A7) showed evidence of association with offspring birth weight at P<5x10-8. The SEM analyses showed at least 7 of the 10 associations were consistent with effects of the maternal genotype acting via the intrauterine environment, rather than via effects of shared alleles with the fetus. Variants, or correlated proxies, at many of the loci had been previously associated with adult traits, including fasting glucose (MTNR1B, GCK and TCF7L2) and sex hormone levels (CYP3A7), and one (EBF1) with gestational duration.\n\nThe identified associations indicate effects of maternal glucose, cytochrome P450 activity and gestational duration, and potential effects of maternal blood pressure and immune function on fetal growth. Further characterization of these associations, for example in mechanistic and causal analyses, will enhance understanding of the potentially modifiable maternal determinants of fetal growth, with the goal of reducing the morbidity and mortality associated with low and high birth weights.

Genetics

Disease variants alter transcription factor levels and methylation of their binding sites

Most disease associated genetic risk factors are non-coding, making it challenging to design experiments to understand their functional consequences1,2. Identification of expression quantitative trait loci (eQTLs) has been a powerful approach to infer downstream effects of disease variants but the large majority remains unexplained.3,4. The analysis of DNA methylation, a key component of the epigenome5, offers highly complementary data on the regulatory potential of genomic regions6,7. However, a large-scale, combined analysis of methylome and transcriptome data to infer downstream effects of disease variants is lacking. Here, we show that disease variants have wide-spread effects on DNA methylation in trans that likely reflect the downstream effects on binding sites of cis-regulated transcription factors. Using data on 3,841 Dutch samples, we detected 272,037 independent cis-meQTLs (FDR < 0.05) and identified 1,907 trait-associated SNPs that affect methylation levels of 10,141 different CpG sites in trans (FDR < 0.05), an eight-fold increase in the number of downstream effects that was known from trans-eQTL studies3,8,9. Trans-meQTL CpG sites are enriched for active regulatory regions, being correlated with gene expression and overlap with Hi-C determined interchromosomal contacts10,11. We detected many trans-meQTL SNPs that affect expression levels of nearby transcription factors (including NFKB1, CTCF and NKX2-3), while the corresponding trans-meQTL CpG sites frequently coincide with its respective binding site. Trans-meQTL mapping therefore provides a strategy for identifying and better understanding downstream functional effects of many disease-associated variants.

Genomics

Hypothesis-free identification of modulators of genetic risk factors

Genetic risk factors often localize in non-coding regions of the genome with unknown effects on disease etiology. Expression quantitative trait loci (eQTLs) help to explain the regulatory mechanisms underlying the association of genetic risk factors with disease. More mechanistic insights can be derived from knowledge of the context, such as cell type or the activity of signaling pathways, influencing the nature and strength of eQTLs. Here, we generated peripheral blood RNA-seq data from 2,116 unrelated Dutch individuals and systematically identified these context-dependent eQTLs using a hypothesis-free strategy that does not require prior knowledge on the identity of the modifiers. Out of the 23,060 significant cis-regulated genes (false discovery rate < 0.05), 2,743 genes (12%) show context-dependent eQTL effects. The majority of those were influenced by cell type composition, revealing eQTLs that are particularly strong in cell types such as CD4+ T-cells, erythrocytes, and even lowly abundant eosinophils. A set of 145 cis-eQTLs were influenced by the activity of the type I interferon signaling pathway and we identified several cis-eQTLs that are modulated by specific transcription factors that bind to the eQTL SNPs. This demonstrates that large-scale eQTL studies in unchallenged individuals can complement perturbation experiments to gain better insight in regulatory networks and their stimuli.

Genetics

Genetic Associations with Subjective Well-Being Also Implicate Depression and Neuroticism

We conducted a genome-wide association study of subjective well-being (SWB) in 298,420 individuals. We also performed auxiliary analyses of depressive symptoms (\"DS\"; N = 161,460) and neuroticism (N = 170,910), both of which have a substantial genetic correlation with SWB [Formula]. We identify three SNPs associated with SWB at genome-wide significance. Two of them are significantly associated with DS in an independent sample. In our auxiliary analyses, we identify 13 additional genome-wide-significant associations: two with DS and eleven with neuroticism, including two inversion polymorphisms. Across our phenotypes, loci regulating expression in central nervous system and adrenal/pancreas tissues are enriched. The discovery of genetic loci associated with the three phenotypes we study has proven elusive; our findings illustrate the payoffs from studying them jointly.\n\nOne Sentence Summary: Using both genome-wide association studies and proxy-phenotype studies, we identify genetic variants associated with subjective well-being, depressive symptoms, and neuroticism.

Genetics

Cell specific eQTL analysis without sorting cells

Expression quantitative trait locus (eQTL) mapping on tissue, organ or whole organism data can detect associations that are generic across cell types. We describe a new method to focus upon specific cell types without first needing to sort cells. We applied the method to whole blood data from 5,683 samples and demonstrate that SNPs associated with Crohn's disease preferentially affect gene expression within neutrophils.

Genetics