Search bioRxivSearch

Biology subjects

Boerwinkle, E.

Publications and source records attributed to Boerwinkle, E..

12 recordsLinked to original sources

Evaluation of the causal effect of fibrinogen on incident coronary heart disease via Mendelian randomization

BackgroundFibrinogen is an essential hemostatic factor and cardiovascular disease risk factor. Early attempts at evaluating the causal effect of fibrinogen on coronary heart disease (CHD) and myocardial infraction (MI) using Mendelian randomization (MR) used single variant approaches, and did not take advantage of recent genome-wide association studies (GWAS) or multi-variant, pleiotropy robust MR methodologies.\n\nMethods and FindingsWe evaluated evidence for a causal effect of fibrinogen on both CHD and MI using MR. We used both an allele score approach and pleiotropy robust MR models. The allele score was composed of 38 fibrinogen-associated variants from recent GWAS. Initial analyses using the allele score incorporated data from 11 European-ancestry prospective cohorts to examine incidence CHD and MI. We also applied 2 sample MR methods with data from a prevalent CHD and MI GWAS. Results are given in terms of the hazard ratio (HR) or odds ratio (OR), depending on the study design, and associated 95% confidence interval (CI).\n\nIn single variant analyses no causal effect of fibrinogen on CHD or MI was observed. In multi-variant analyses using incidence CHD cases and the allele score approach, the estimated causal effect (HR) of a 1 g/L higher fibrinogen concentration was 1.62 (CI = 1.12, 2.36) when using incident cases and the allele score approach. In 2 sample MR analyses that accounted for pleiotropy, the causal estimate (OR) was reduced to 1.18 (CI = 0.98, 1.42) and 1.09 (CI = 0.89, 1.33) in the 2 most precise (smallest CI) models, out of 4 models evaluated. In the 2 sample MR analyses for MI, there was only very weak evidence of a causal effect in only 1 out of 4 models.\n\nConclusionsA small causal effect of fibrinogen on CHD is observed using multi-variant MR approaches which account for pleiotropy, but not single variant MR approaches. Taken together, results indicate that even with large sample sizes and multi-variant approaches MR analyses still cannot exclude the null when estimating the causal effect of fibrinogen on CHD, but that any potential causal effect is likely to be much smaller than observed in epidemiological studies.\n\nAuthor SummaryInitial Mendelian Randomization (MR) analyses of the causal effect of fibrinogen on coronary heart disease (CHD) utilized single variants and did not take advantage of modern, multivariant approaches. This manuscript provides an important update to these initial analyses by incorporating larger sample sizes and employing multiple, modern multi-variant MR approaches to account for pleiotropy. We used incident cases to perform a MR study of the causal effect of fibrinogen on incident CHD and the nested outcome of myocardial infarction (MI) using an allele score approach. Then using data from a case-control genome-wide association study for CHD and MI we performed two sample MR analyses with multiple, pleiotropy robust approaches. Overall, the results indicated that associations between fibrinogen and CHD in observational studies are likely upwardly biased from any underlying causal effect. Single variant MR approaches show little evidence of a causal effect of fibrinogen on CHD or MI. Multi-variant MR analyses of fibrinogen on CHD indicate there may be a small positive effect, however this result needs to be interpreted carefully as the 95% confidence intervals were still consistent with a null effect. Multi-variant MR approaches did not suggest evidence of even a small causal effect of fibrinogen on MI.

genetics

Atlas-CNV: a validated approach to call Single-Exon CNVs in the eMERGESeq gene panel

PurposeTo provide a validated method to confidently identify exon-containing copy number variants (CNVs), with a low false discovery rate (FDR), in targeted sequencing data from a clinical laboratory with particular focus on single-exon CNVs.\n\nMethodsDNA sequence coverage data are normalized within each sample and subsequently exonic CNVs are identified in a batch of samples (midpool), when the target log2 ratio of the sample to the batch median exceeds defined thresholds. The quality of exonic CNV calls is assessed by C-scores (Z-like scores) using thresholds derived from gold standard samples and simulation studies. We integrate an ExonQC threshold to lower FDR and compare performance with alternate software (VisCap).\n\nResultsThirteen CNVs were used as a truth set to validate Atlas-CNV and compared with VisCap. We demonstrated FDR reduction in validation, simulation and 10,926 eMERGESeq samples without sensitivity loss. Sixty-four multi-exon and 29 single-exon CNVs with high C-scores were assessed by MLPA.\n\nConclusionsAtlas-CNV is validated as a method to identify exonic CNVs in targeted sequencing data generated in the clinical laboratory. The ExonQC and C-score assignment can reduce FDR (identification of targets with high variance) and improve calling accuracy of single-exon CNVs respectively. We proposed guidelines and criteria to identify high confidence single-exon CNVs.

genomics

Parliament2: Fast Structural Variant Calling Using Optimized Combinations of Callers

Here we present Parliament2 - a structural variant caller which combines multiple best-in-class structural variant callers to create a highly accurate callset. This captures more events than the individual callers achieve independently. Parliament2 uses a call-overlap-genotype approach that is highly extensible to new methods and presents users the choice to run some or all of Breakdancer, Breakseq, CNVnator, Delly, Lumpy, and Manta to run. Parliament2 applies an additional parallelization framework to speed certain callers and executes these in parallel, taking advantage of the different resource requirements to complete structural variant calling much faster than running the programs individually. Parliament2 is available as a Docker container, which pre-installs all required dependencies. This allows users to run any caller with easy installation and execution. This Docker container can easily be deployed in cloud or local environments and is available as an app on DNAnexus.

bioinformatics

Plasma metabolomics and incidence of atrial fibrillation: the Atherosclerosis Risk in Communities (ARIC) Study

We have previously identified associations of two circulating secondary bile acids (glycocholenate and glycolithocolate sulfate) with atrial fibrillation (AF) risk among blacks. We aimed to replicate these findings in an independent sample including both whites and blacks, and performed a new metabolomic analysis in the combined sample. We studied 3,922 participants from the ARIC cohort followed between 1987 and 2013. Of these, 1,919 had been included in the prior analysis and 2,003 were new samples. Metabolomic profiling was done in baseline serum samples using gas and liquid chromatography mass spectrometry. AF was ascertained from electrocardiograms, hospitalizations, and death certificates. We used multivariable Cox regression to estimate hazard ratios (HR) and 95% confidence intervals (95%CI) of AF by one standard deviation difference of metabolite levels. Over a mean follow-up of 20 years, 608 participants developed AF. Glycocholenate sulfate was associated with AF in the replication and combined samples (HR 1.10, 95%CI 1.00, 1.21 and HR 1.13, 95%CI 1.04, 1.22, respectively). Glycolithocolate sulfate was not related to AF risk in the replication sample (HR 1.02, 95%CI 0.92, 1.13). An analysis of 245 metabolites in the combined cohort identified three additional metabolites associated with AF after multiple-comparison correction: pseudouridine (HR 1.18, 95%CI 1.10, 1.28), uridine (HR 0.86, 95%CI 0.79, 0.93) and acisoga (HR 1.17, 95%CI 1.09, 1.26). To conclude, we replicated a prospective association between a previously identified secondary bile acid, glycocholenate sulfate, and AF incidence, and identified new metabolites involved in nucleoside and polyamine metabolism as markers of AF risk.

epidemiology

PROTEIN-CODING VARIANTS IMPLICATE NOVEL GENES RELATED TO LIPID HOMEOSTASIS CONTRIBUTING TO BODY FAT DISTRIBUTION

Body fat distribution is a heritable risk factor for a range of adverse health consequences, including hyperlipidemia and type 2 diabetes. To identify protein-coding variants associated with body fat distribution, assessed by waist-to-hip ratio adjusted for body mass index, we analyzed 228,985 predicted coding and splice site variants available on exome arrays in up to 344,369 individuals from five major ancestries for discovery and 132,177 independent European-ancestry individuals for validation. We identified 15 common (minor allele frequency, MAF[&ge;]5%) and 9 low frequency or rare (MAF<5%) coding variants that have not been reported previously. Pathway/gene set enrichment analyses of all associated variants highlight lipid particle, adiponectin level, abnormal white adipose tissue physiology, and bone development and morphology as processes affecting fat distribution and body shape. Furthermore, the cross-trait associations and the analyses of variant and gene function highlight a strong connection to lipids, cardiovascular traits, and type 2 diabetes. In functional follow-up analyses, specifically in Drosophila RNAi-knockdown crosses, we observed a significant increase in the total body triglyceride levels for two genes (DNAH10 and PLXND1). By examining variants often poorly tagged or entirely missed by genome-wide association studies, we implicate novel genes in fat distribution, stressing the importance of interrogating low-frequency and protein-coding variants.

genetics

Quality Control and Integration of Genotypes from Two Calling Pipelines for Whole Genome Sequence Data in the Alzheimer’s Disease Sequencing Project

The Alzheimers Disease Sequencing Project (ADSP) performed whole genome sequencing (WGS) of 584 subjects from 111 multiplex families at three sequencing centers. Genotype calling of single nucleotide variants (SNVs) and insertion-deletion variants (indels) was performed centrally using GATK-HaplotypeCaller and Atlas V2. The ADSP Quality Control (QC) Working Group applied QC protocols to project-level variant call format files (VCFs) from each pipeline, and developed and implemented a novel protocol, termed \"consensus calling,\" to combine genotype calls from both pipelines into a single high-quality set. QC was applied to autosomal bi-allelic SNVs and indels, and included pipeline-recommended QC filters, variant-level QC, and sample-level QC. Low-quality variants or genotypes were excluded, and sample outliers were noted. Quality was assessed by examining Mendelian inconsistencies (MIs) among 67 parent-offspring pairs, and MIs were used to establish additional genotype-specific filters for GATK calls. After QC, 578 subjects remained. Pipeline-specific QC excluded ~12.0% of GATK and 14.5% of Atlas SNVs. Between pipelines, ~91% of SNV genotypes across all QCed variants were concordant; 4.23% and 4.56% of genotypes were exclusive to Atlas or GATK, respectively; the remaining ~0.01% of discordant genotypes were excluded. For indels, variant-level QC excluded ~36.8% of GATK and 35.3% of Atlas indels. Between pipelines, ~55.6% of indel genotypes were concordant; while 10.3% and 28.3% were exclusive to Atlas or GATK, respectively; and ~0.29% of discordant genotypes were. The final WGS consensus dataset contains 27,896,774 SNVs and 3,133,926 indels and is publicly available.\n\nAbbreviationsAD, Alzheimers disease; QC, Quality Control; LSSAC, Large-Scale Sequencing and Analysis Center; Broad, Broad Institute Genomics Service; Baylor, Baylor College of Medicine Human Genome Sequencing Center; WashU, Washington University-St. Louis McDonnell Genome Institute; WGS, whole genome sequencing; WES, whole exome sequencing; indel, insertion-deletion variants; VCF, variant control format; MI, Mendelian inconsistency; MC, Mendelian consistency; GWAS, genome-wide association study; VR, referent allele read depth; DP, overall read depth; MS, mapping score; GQ, genotype quality score; Ti/Tv, Transition/Transversion; CS, concordance code

genetics

xAtlas: Scalable small variant calling across heterogeneous next-generation sequencing experiments

MotivationThe rapid development of next-generation sequencing (NGS) technologies has lowered the barriers to genomic data generation, resulting in millions of samples sequenced across diverse experimental designs. The growing volume and heterogeneity of these sequencing data complicate the further optimization of methods for identifying DNA variation, especially considering that curated highconfidence variant call sets commonly used to evaluate these methods are generally developed by reference to results from the analysis of comparatively small and homogeneous sample sets.\n\nResultsWe have developed xAtlas, an application for the identification of single nucleotide variants (SNV) and small insertions and deletions (indels) in NGS data. xAtlas is easily scalable and enables execution and retraining with rapid development cycles. Generation of variant calls in VCF or gVCF format from BAM or CRAM alignments is accomplished in less than one CPU-hour per 30x short-read human whole-genome. The retraining capabilities of xAtlas allow its core variant evaluation models to be optimized on new sample data and user-defined truth sets. Obtaining SNV and indels calls from xAtlas can be achieved more than 40 times faster than established methods while retaining the same accuracy.\n\nAvailabilityFreely available under a BSD 3-clause license at https://github.com/jfarek/xatlas.\n\nContactfarek@bcm.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Genome-wide association meta-analysis of PR interval identifies 47 novel loci associated with atrial and atrioventricular electrical activity

Electrocardiographic PR interval measures atrial and atrioventricular depolarization and conduction, and abnormal PR interval is a risk factor for atrial fibrillation and heart block. We performed a genome-wide association study in over 92,000 individuals of European descent and identified 44 loci associated with PR interval (34 novel). Examination of the 44 loci revealed known and novel biological processes involved in cardiac atrial electrical activity, and genes in these loci were highly over-represented in several cardiac disease processes. Nearly half of the 61 independent index variants in the 44 loci were associated with atrial or blood transcript expression levels, or were in high linkage disequilibrium with one or more missense variants. Cardiac regulatory regions of the genome as measured by cardiac DNA hypersensitivity sites were enriched for variants associated with PR interval, compared to non-cardiac regulatory regions. Joint analyses combining PR interval with heart rate, QRS interval, and atrial fibrillation identified additional new pleiotropic loci. The majority of associations discovered in European-descent populations were also present in African-American populations. Meta-analysis examining over 105,000 individuals of African and European descent identified additional novel PR loci. These additional analyses identified another 13 novel loci. Together, these findings underscore the power of GWAS to extend knowledge of the molecular underpinnings of clinical processes.

genetics

Genome-wide Association Study Links APOEϵ4 and BACE1 Variants with Plasma Amyloid β Levels

INTRODUCTIONThere is increasing interest in plasma A{beta} as an endophenotype and biomarker of Alzheimers disease (AD). Identifying the genetic determinants of plasma A{beta} levels may elucidate important processes that determine plasma A{beta} measures. METHODSWe included 12,369 non-demented participants derived from eight population-based studies. Imputed genetic data and plasma A{beta}1-40, A{beta}1-42 levels and A{beta}1-42/A{beta}1-40 ratio were used to perform genome-wide association studies, gene-based and pathway analyses. Significant variants and genes were followed-up for the association with PET A{beta} deposition and AD risk. RESULTSSingle-variant analysis identified associations across APOE for A{beta}1-42 and A{beta}1-42/A{beta}1-40 ratio, and BACE1 for A{beta}1-40. Gene-based analysis of A{beta}1-40 additionally identified associations for APP, PSEN2, CCK and ZNF397. There was suggestive interaction between a BACE1 variant and APOE{varepsilon}4 on brain A{beta} deposition. DISCUSSIONIdentification of variants near/in known major A{beta}-processing genes strengthens the relevance of plasma-A{beta} levels both as an endophenotype and a biomarker of AD.

genetics

Genetic Diversity Turns a New PAGE in Our Understanding of Complex Traits

Summary/AbstractGenome-wide association studies (GWAS) have laid the foundation for investigations into the biology of complex traits, drug development, and clinical guidelines. However, the dominance of European-ancestry populations in GWAS creates a biased view of the role of human variation in disease, and hinders the equitable translation of genetic associations into clinical and public health applications. The Population Architecture using Genomics and Epidemiology (PAGE) study conducted a GWAS of 26 clinical and behavioral phenotypes in 49,839 non-European individuals. Using strategies designed for analysis of multi-ethnic and admixed populations, we confirm 574 GWAS catalog variants across these traits, and find 38 secondary signals in known loci and 27 novel loci. Our data shows strong evidence of effect-size heterogeneity across ancestries for published GWAS associations, substantial benefits for fine-mapping using diverse cohorts, and insights into clinical implications. We strongly advocate for continued, large genome-wide efforts in diverse populations to reduce health disparities.

genetics

Association between Mitochondrial DNA Copy Number and Sudden Cardiac Death: Findings from the Atherosclerosis Risk in Communities Study (ARIC)

AimsSudden cardiac death (SCD) is a major public health burden. Mitochondrial dysfunction has been implicated in a wide range of cardiovascular diseases including cardiomyopathy, heart failure, and arrhythmias, but it is unknown if it also contributes to SCD risk. We sought to examine the prospective association between mtDNA copy number (mtDNA-CN), a surrogate marker of mitochondrial function, and SCD risk.\n\nMethods and ResultsWe measured baseline mtDNA-CN in 11,093 participants from the Atherosclerosis Risk in Communities (ARIC) study. mtDNA-CN was calculated from probe intensities of mitochondrial single nucleotide polymorphisms (SNP) on the Affymetrix Genome-Wide Human SNP Array 6.0. SCD was defined as a sudden pulseless condition presumed due to a ventricular tachyarrhythmia in a previously stable individual without evidence of a non-cardiac cause of cardiac arrest. SCD cases were reviewed and adjudicated by an expert committee. During a median follow-up of 20.4 years, we observed 361 SCD cases. After adjusting for age, race, sex, and center, the hazard ratio (HR) for SCD comparing the 1st to the 5th quintiles of mtDNA-CN was 2.24 (95% CI 1.58 to 3.19; p-trend <0.001). When further adjusting for traditional CVD risk factors, prevalent CHD, heart rate, and QT interval duration, the association remained statistically significant. Spline regression models showed that the association was approximately linear over the range of mtDNA-CN values. No apparent interaction by race or by sex was detected.\n\nConclusionIn this community-based prospective study, mtDNA-CN in peripheral blood was inversely associated with the risk of SCD.

epidemiology

Hardy Weinberg Exact Test In Large Scale Variant Calling Quality Control

Hardy Weinberg Equilibrium (HWE) test is widely used as a quality control measure to detect sequencing artifacts like mismapping, allelic dropout and biases. However, in the high throughput sequencing era, where the sample size is beyond a thousand scale, the utility of HWE test in reducing the false positive rate remains unclear. In this paper, we demonstrate that HWE test has limited power in identifying sequencing artifacts when the variant allele frequency is lower than 1% in a variant call set produced from more than five thousand whole genome sequenced samples from two homogeneous populations. We develop a novel strategy of implementing HWE filtering in which we incorporate site frequency spectrum information and determine the p-value cutoff which optimizes the tradeoff between sensitivity and specificity. The novel strategy is shown to outperform the exact test of HWE with an empirical constant p-value cutoff regardless of the sequencing sample size. We also present best practice recommendations for identifying possible sources of false positives from large sequencing datasets based on an analysis of intrinsic biases in the variant calling process. Our novel strategy of determining the HWE test p-value cutoff and applying the test to the common variants provides a practical approach for the variant level quality controls in the upcoming sequencing projects with tens to hundreds of thousand of samples.

bioinformatics